All topics
Back to DailyYouTubeEpisode 5 of 9 · CS329A Self-Improving AI Agents: Stanford's Complete CourseDuration:1:14:55
CS329A Self-Improving AI Agents, Part 5: Planning and Multi-Step Reasoning
Stanford CS329A Self-Improving AI Agents | Part 5 | Planning and Multi-Step Reasoning
SOStanford Online@stanfordonlineFull transcript
English
Summary:LATS brings Monte Carlo Tree Search to language agents with LLM-judge scoring; SPRINT fine-tunes reasoners to emit independent plans that run in parallel; SWiRL trains on synthetic multi-step tool-use trajectories and generalizes without executing tools during training.
Core points (3)
Core points (3)
- 1LATS combines reasoning, acting, and MCTS with LLM-judge and self-consistency scoring.
- 2SPRINT cuts sequential tokens by generating independent plans for parallel execution.
- 3SWiRL trains on synthetic tool trajectories without calling tools during training.