CS329A Self-Improving AI Agents, Part 3: Robust Verification

Stanford CS329A Self-Improving AI Agents | Part 3 | Robust Verification

SOStanford Online@stanfordonline

Full transcript

English

Summary:Four papers trace verification: GSM8K with outcome-based verifiers, process supervision on PRM800K, Math-Shepherd's automated step labels, and Weaver's ensembles of weak verifiers. Credit assignment and reward hacking recur as the failure modes that decide whether verification can be trusted.

Watch on YouTube
Core points (3)

Core points (3)

  1. 1Process supervision assigns credit per step; outcome supervision judges only the final answer.
  2. 2Math-Shepherd removes human labeling by bootstrapping step correctness from rollouts.
  3. 3Ensembles of weak verifiers approach the reliability of a single strong judge.