All topics
Back to DailyYouTubeEpisode 7 of 9 · CS329A Self-Improving AI Agents: Stanford's Complete CourseDuration:1:12:26
CS329A Self-Improving AI Agents, Part 7: Self-Improvement and Deep Research Agents
Stanford CS329A Self-Improving AI Agents | Part 7 | Self-Improvement and Deep Research Agents
SOStanford Online@stanfordonlineFull transcript
English
Summary:AlphaCode reached competitive-programming results via huge sample budgets and clustering; AlphaCode2 fine-tuned Gemini Pro with a learned scoring model to reach the 85th percentile. Search-O1 applies the same lesson to research: retrieve at uncertainty points and reason over documents.
Core points (3)
Core points (3)
- 1Solve rate scales with sample budget, but selection and clustering become the bottleneck.
- 2Search-O1 fires retrieval only when the reasoner signals uncertainty.
- 3Prompted search (Search-O1) and RL-trained search (Search-R1) trade flexibility for stability.