CS329A Self-Improving AI Agents, Part 5: Planning and Multi-Step Reasoning

Stanford CS329A Self-Improving AI Agents | Part 5 | Planning and Multi-Step Reasoning

SOStanford Online@stanfordonline

Full transcript

English

Summary:LATS brings Monte Carlo Tree Search to language agents with LLM-judge scoring; SPRINT fine-tunes reasoners to emit independent plans that run in parallel; SWiRL trains on synthetic multi-step tool-use trajectories and generalizes without executing tools during training.

Watch on YouTube
Core points (3)

Core points (3)

  1. 1LATS combines reasoning, acting, and MCTS with LLM-judge and self-consistency scoring.
  2. 2SPRINT cuts sequential tokens by generating independent plans for parallel execution.
  3. 3SWiRL trains on synthetic tool trajectories without calling tools during training.