CS329A Self-Improving AI Agents, Part 2: Test-Time Compute Scaling

Stanford CS329A Self-Improving AI Agents | Part 2 | Test-Time Compute Scaling

SOStanford Online@stanfordonline

Full transcript

English

Summary:Solve rates follow a power law as parallel samples grow, driven by a long tail of rarely solved problems. The lecture separates majority voting from oracle verification to define the generation-verification gap, then compares parallel sampling with sequential revision guided by reward models.

Watch on YouTube
Core points (3)

Core points (3)

  1. 1Solve rate follows a power law in the number of parallel samples per problem.
  2. 2The generation-verification gap measures how far majority voting sits from oracle selection.
  3. 3Process and outcome reward models guide test-time search better than blind resampling.