All topics
Back to DailyYouTubeEpisode 4 of 9 · CS329A Self-Improving AI Agents: Stanford's Complete CourseDuration:1:11:13
CS329A Self-Improving AI Agents, Part 4: Learning from Feedback with Tools and Code
Stanford CS329A Self-Improving AI Agents | Part 4 | Learning from Feedback with Tools/Code
SOStanford Online@stanfordonlineFull transcript
English
Summary:Three papers show agents learning from non-human feedback: ReAct interleaves reasoning with tool calls, RLEF trains coding agents on unit-test results inside PPO, and Constitutional AI replaces human preference labels with principle-guided self-critique and reinforcement learning from AI feedback.
Core points (3)
Core points (3)
- 1ReAct alternates chain-of-thought reasoning with tool actions on HotpotQA and WebShop.
- 2RLEF shows execution feedback from unit tests is a scalable reward for coding agents.
- 3Constitutional AI swaps human raters for written principles and self-critique.