CS329A Self-Improving AI Agents, Part 7: Self-Improvement and Deep Research Agents

Stanford CS329A Self-Improving AI Agents | Part 7 | Self-Improvement and Deep Research Agents

SOStanford Online@stanfordonline

Full transcript

English

Summary:AlphaCode reached competitive-programming results via huge sample budgets and clustering; AlphaCode2 fine-tuned Gemini Pro with a learned scoring model to reach the 85th percentile. Search-O1 applies the same lesson to research: retrieve at uncertainty points and reason over documents.

Watch on YouTube
Core points (3)

Core points (3)

  1. 1Solve rate scales with sample budget, but selection and clustering become the bottleneck.
  2. 2Search-O1 fires retrieval only when the reasoner signals uncertainty.
  3. 3Prompted search (Search-O1) and RL-trained search (Search-R1) trade flexibility for stability.