CS329A Self-Improving AI Agents, Part 4: Learning from Feedback with Tools and Code

Stanford CS329A Self-Improving AI Agents | Part 4 | Learning from Feedback with Tools/Code

SOStanford Online@stanfordonline

Full transcript

English

Summary:Three papers show agents learning from non-human feedback: ReAct interleaves reasoning with tool calls, RLEF trains coding agents on unit-test results inside PPO, and Constitutional AI replaces human preference labels with principle-guided self-critique and reinforcement learning from AI feedback.

Watch on YouTube
Core points (3)

Core points (3)

  1. 1ReAct alternates chain-of-thought reasoning with tool actions on HotpotQA and WebShop.
  2. 2RLEF shows execution feedback from unit tests is a scalable reward for coding agents.
  3. 3Constitutional AI swaps human raters for written principles and self-critique.