All topics
Back to DailyYouTubeEpisode 2 of 9 · CS329A Self-Improving AI Agents: Stanford's Complete CourseDuration:1:03:20
CS329A Self-Improving AI Agents, Part 2: Test-Time Compute Scaling
Stanford CS329A Self-Improving AI Agents | Part 2 | Test-Time Compute Scaling
SOStanford Online@stanfordonlineFull transcript
English
Summary:Solve rates follow a power law as parallel samples grow, driven by a long tail of rarely solved problems. The lecture separates majority voting from oracle verification to define the generation-verification gap, then compares parallel sampling with sequential revision guided by reward models.
Core points (3)
Core points (3)
- 1Solve rate follows a power law in the number of parallel samples per problem.
- 2The generation-verification gap measures how far majority voting sits from oracle selection.
- 3Process and outcome reward models guide test-time search better than blind resampling.