All topics
Back to DailyYouTubeEpisode 6 of 9 · CS329A Self-Improving AI Agents: Stanford's Complete CourseDuration:1:12:38
CS329A Self-Improving AI Agents, Part 6: Train-Time Scaling and Scaling RL
Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL
SOStanford Online@stanfordonlineFull transcript
English
Summary:STaR bootstraps reasoning chains by rationalization and filtering on correctness; DeepSeekMath contributes GRPO, a memory-efficient alternative to PPO; DAPO fixes entropy collapse with asymmetric clipping and dynamic sampling. On AIME, smaller models begin matching much larger ones.
Core points (3)
Core points (3)
- 1STaR grows new reasoning chains by rationalizing the problems it first got wrong.
- 2GRPO replaces PPO's critic with group-relative advantages to cut memory cost.
- 3DAPO's asymmetric clipping and dynamic sampling prevent entropy collapse in long RL runs.