All topics
Back to DailyYouTubeEpisode 8 of 9 · CS329A Self-Improving AI Agents: Stanford's Complete CourseDuration:1:15:17
CS329A Self-Improving AI Agents, Part 8: Agentic Evaluations and Long-Horizon Tasks
Stanford CS329A Self-Improving AI Agents | Part 8 | Agentic Evaluations and Long Horizon Tasks
SOStanford Online@stanfordonlineFull transcript
English
Summary:METR measures the task duration models complete at 50 and 80 percent reliability; capability has doubled roughly every seven months. GDPval scores model work against professionals in 44 occupations, Deep Scholar Bench tests literature reviews, and recurring failure modes close the lecture.
Core points (3)
Core points (3)
- 1METR shows model time horizons double roughly every seven months.
- 2GDPval win rates climbed from 12.4 to 47.6 percent against working professionals.
- 3Typical agent failures are planning, tool choice, premature abandonment, and loops.