This episode examines Guided Policy Search, a 2013 method from Sergey Levine and Vladlen Koltun that lets flexible neural-network policies control robots without falling into the poor local optima that plague direct policy search over high-dimensional parameter spaces. The discussion traces the paper's teacher-student structure: differential dynamic programming (DDP), a model-based trajectory optimizer rooted in 1960s optimal control theory, generates high-reward example trajectories for specific starting conditions, and the neural-network student learns to match and generalize this behavior via policy gradients and importance sampling rather than naive imitation. A key distinction drawn out is why this differs from imitation learning approaches like DAGGER — DDP's guidance is only locally valid, so the method needs an objective built to maximize reward everywhere, not just mimic a narrow expert trajectory. The conversation connects DDP's backward pass to Bellman recursion and the broader LQR/Kalman-filter lineage, and explains how importance sampling lets the same batch of guiding samples be reused across many gradient steps, which matters when real-hardware data collection is expensive. Listeners interested in the historical roots of modern reinforcement learning — and how classical control theory was fused with neural networks years before this became standard practice — will find the episode's walkthrough of the underlying mechanics clarifying.
Sources:
1. Guided Policy Search: Teaching Neural Nets via Trajectory Optimization
https://proceedings.mlr.press/v28/levine13.pdf2. Learning Neural Network Policies with Guided Policy Search under Unknown Dynamics — Sergey Levine, Pieter Abbeel, 2014
https://scholar.google.com/scholar?q=Learning+Neural+Network+Policies+with+Guided+Policy+Search+under+Unknown+Dynamics3. End-to-End Training of Deep Visuomotor Policies — Sergey Levine, Chelsea Finn, Trevor Darrell, Pieter Abbeel, 2016
https://scholar.google.com/scholar?q=End-to-End+Training+of+Deep+Visuomotor+Policies4. Trust Region Policy Optimization — John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, Pieter Abbeel, 2015
https://scholar.google.com/scholar?q=Trust+Region+Policy+Optimization5. Reinforcement Learning of Motor Skills with Policy Gradients — Jan Peters, Stefan Schaal, 2008
https://scholar.google.com/scholar?q=Reinforcement+Learning+of+Motor+Skills+with+Policy+Gradients6. Differential Dynamic Programming — David Jacobson, David Mayne, 1970
https://scholar.google.com/scholar?q=Differential+Dynamic+Programming7. A Generalized Iterative LQG Method for Locally-Optimal Feedback Control of Constrained Nonlinear Stochastic Systems — Emanuel Todorov, Weiwei Li, 2005
https://scholar.google.com/scholar?q=A+Generalized+Iterative+LQG+Method+for+Locally-Optimal+Feedback+Control+of+Constrained+Nonlinear+Stochastic+Systems8. Synthesis and Stabilization of Complex Behaviors through Online Trajectory Optimization — Yuval Tassa, Tom Erez, Emanuel Todorov, 2012
https://scholar.google.com/scholar?q=Synthesis+and+Stabilization+of+Complex+Behaviors+through+Online+Trajectory+Optimization9. Aggressive Driving with Model Predictive Path Integral Control — Grady Williams, Paul Drews, Brian Goldfain, James Rehg, Evangelos Theodorou, 2016
https://scholar.google.com/scholar?q=Aggressive+Driving+with+Model+Predictive+Path+Integral+Control10. Eligibility Traces for Off-Policy Policy Evaluation — Doina Precup, Richard Sutton, Satinder Singh, 2000
https://scholar.google.com/scholar?q=Eligibility+Traces+for+Off-Policy+Policy+Evaluation11. Learning from Scarce Experience — Leonid Peshkin, Christian Shelton, 2002
https://scholar.google.com/scholar?q=Learning+from+Scarce+Experience12. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms13. PILCO: A Model-Based and Data-Efficient Approach to Policy Search — Deisenroth, M. and Rasmussen, C., 2011
https://scholar.google.com/scholar?q=PILCO%3A+A+Model-Based+and+Data-Efficient+Approach+to+Policy+Search14. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAGGER) — Ross, S., Gordon, G., and Bagnell, A., 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAGGER%2915. On a Connection Between Importance Sampling and the Likelihood Ratio Policy Gradient — Tang, J. and Abbeel, P., 2010
https://scholar.google.com/scholar?q=On+a+Connection+Between+Importance+Sampling+and+the+Likelihood+Ratio+Policy+Gradient16. Approximately Optimal Approximate Reinforcement Learning — Kakade, S. and Langford, J., 2002
https://scholar.google.com/scholar?q=Approximately+Optimal+Approximate+Reinforcement+Learning17. SIMBICON: Simple Biped Locomotion Control — Yin, K., Loken, K., and van de Panne, M., 2007
https://scholar.google.com/scholar?q=SIMBICON%3A+Simple+Biped+Locomotion+ControlInteractive Visualization: Guided Policy Search: Teaching Neural Nets via Trajectory Optimization