ai研究
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| MIT proposes VPO: Vectorized rewards replace scalars to maintain diversity in LLM test-time search |
|
0 | 14 | May 23, 2026 |
| Former DeepMind VP Nando de Freitas: Pure imitation learning can lead to reward-maximizing behavior without needing handcrafted reward functions |
|
0 | 7 | May 22, 2026 |