Stanford CS329A course playlist cover

YouTube series9 episodes10h 40m

Stanford Online

CS329A Self-Improving AI Agents: Stanford's Complete Course

CS329A Self-Improving AI Agents

Core points (3)

  • The generation-verification gap is the course's spine: sampling is cheap, checking is not.
  • Self-improvement loops need process supervision, robust verifiers, and reward-hacking defenses.
  • METR's doubling clock and GDPval put agent progress on a measurable ladder.
Start with episode 1

Episodes

  1. Stanford CS329A Part 1 lecture thumbnail

    01 · 1h 9m

    CS329A Self-Improving AI Agents, Part 1: Course Overview

    Part 1 frames the course: from scaling laws to agent workflows.

    Read
  2. Stanford CS329A Part 2 lecture thumbnail

    02 · 1h 3m

    CS329A Self-Improving AI Agents, Part 2: Test-Time Compute Scaling

    Part 2 quantifies test-time scaling and names the generation-verification gap.

    Read
  3. Stanford CS329A Part 3 lecture thumbnail

    03 · 1h 12m

    CS329A Self-Improving AI Agents, Part 3: Robust Verification

    Part 3 digs into verifiers: process vs outcome supervision, reward hacking.

    Read
  4. Stanford CS329A Part 4 lecture thumbnail

    04 · 1h 11m

    CS329A Self-Improving AI Agents, Part 4: Learning from Feedback with Tools and Code

    Part 4 shows agents learning from tool actions, unit tests, and AI critique.

    Read
  5. Stanford CS329A Part 5 lecture thumbnail

    05 · 1h 14m

    CS329A Self-Improving AI Agents, Part 5: Planning and Multi-Step Reasoning

    Part 5 adds planning: tree search, parallel plans, and synthetic trajectories.

    Read
  6. Stanford CS329A Part 6 lecture thumbnail

    06 · 1h 12m

    CS329A Self-Improving AI Agents, Part 6: Train-Time Scaling and Scaling RL

    Part 6 moves the scaling inside training: STaR, GRPO, and DAPO on AIME.

    Read
  7. Stanford CS329A Part 7 lecture thumbnail

    07 · 1h 12m

    CS329A Self-Improving AI Agents, Part 7: Self-Improvement and Deep Research Agents

    Part 7 scales self-improvement: AlphaCode's budget, Search-O1's research loop.

    Read
  8. Stanford CS329A Part 8 lecture thumbnail

    08 · 1h 15m

    CS329A Self-Improving AI Agents, Part 8: Agentic Evaluations and Long-Horizon Tasks

    Part 8 measures agents honestly: time horizons, GDPval, and failure modes.

    Read
  9. Stanford CS329A Part 9 lecture thumbnail

    09 · 1h 7m

    CS329A Self-Improving AI Agents, Part 9: Future Research Areas

    Part 9 closes on open problems: meta-verification, energy, continual learning.

    Read