CS329A Self-Improving AI Agents, Part 6: Train-Time Scaling and Scaling RL

Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL

SOStanford Online@stanfordonline

Full transcript

English

Summary:STaR bootstraps reasoning chains by rationalization and filtering on correctness; DeepSeekMath contributes GRPO, a memory-efficient alternative to PPO; DAPO fixes entropy collapse with asymmetric clipping and dynamic sampling. On AIME, smaller models begin matching much larger ones.

Watch on YouTube
Core points (3)

Core points (3)

  1. 1STaR grows new reasoning chains by rationalizing the problems it first got wrong.
  2. 2GRPO replaces PPO's critic with group-relative advantages to cut memory cost.
  3. 3DAPO's asymmetric clipping and dynamic sampling prevent entropy collapse in long RL runs.