← All episodes An Algorithmic Perspective on Imitation Learning

An Algorithmic Perspective on Imitation Learning

Sep 12, 2026
This episode examines a 2018 survey, "An Algorithmic Perspective on Imitation Learning," which frames robot skill acquisition as an alternative to brittle manual programming or fragile reward engineering. It contrasts two core approaches: behavioral cloning, which treats the problem as supervised learning but suffers from compounding errors when the policy drifts into states the expert never demonstrated, and inverse reinforcement learning, which recovers the expert's underlying reward function before solving for a policy, trading computational cost for better generalization. Concrete examples like the ALVINN self-driving system, AlphaGo's use of expert-game pretraining, and Dynamic Movement Primitives illustrate how these ideas played out in practice, with DMPs offered as a hand-structured counterpoint to fully learned neural approaches. Listeners interested in the tradeoffs between hand-designed structure and end-to-end learning, or in how robotics tackled these problems just before deep learning reshaped the field, will find the historical framing useful for understanding today's imitation-learning methods.
Sources:
1. An Algorithmic Perspective on Imitation Learning — Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J. Andrew Bagnell, Pieter Abbeel, Jan Peters, 2018
http://arxiv.org/abs/1811.06711
2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29
3. End to End Learning for Self-Driving Cars — Mariusz Bojarski, et al. (NVIDIA), 2016
https://scholar.google.com/scholar?q=End+to+End+Learning+for+Self-Driving+Cars
4. Generative Adversarial Imitation Learning (GAIL) — Jonathan Ho, Stefano Ermon, 2016
https://scholar.google.com/scholar?q=Generative+Adversarial+Imitation+Learning+%28GAIL%29
5. ALVINN: An Autonomous Land Vehicle in a Neural Network — Dean Pomerleau, 1989
https://scholar.google.com/scholar?q=ALVINN%3A+An+Autonomous+Land+Vehicle+in+a+Neural+Network
6. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion — Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, Shuran Song, 2023
https://scholar.google.com/scholar?q=Diffusion+Policy%3A+Visuomotor+Policy+Learning+via+Action+Diffusion
7. Algorithms for Inverse Reinforcement Learning — Andrew Ng, Stuart Russell, 2000
https://scholar.google.com/scholar?q=Algorithms+for+Inverse+Reinforcement+Learning
8. Maximum Entropy Inverse Reinforcement Learning — Brian Ziebart, Andrew Maas, J. Andrew Bagnell, Anind Dey, 2008
https://scholar.google.com/scholar?q=Maximum+Entropy+Inverse+Reinforcement+Learning
9. Deep Reinforcement Learning from Human Preferences — Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, Dario Amodei, 2017
https://scholar.google.com/scholar?q=Deep+Reinforcement+Learning+from+Human+Preferences
10. Dynamical Movement Primitives: Learning Attractor Models for Motor Behaviors — Auke Ijspeert, Jun Nakanishi, Heiko Hoffmann, Peter Pastor, Stefan Schaal, 2013 (Neural Computation; building on the authors' earlier 2002/2003 conference papers)
https://scholar.google.com/scholar?q=Dynamical+Movement+Primitives%3A+Learning+Attractor+Models+for+Motor+Behaviors
11. Probabilistic Movement Primitives — Alexandros Paraschos, Christian Daniel, Jan Peters, Gerhard Neumann, 2013
https://scholar.google.com/scholar?q=Probabilistic+Movement+Primitives
12. Learning and Generalization of Motor Skills by Learning from Demonstration — Peter Pastor, Heiko Hoffmann, Tamim Asfour, Stefan Schaal, 2009
https://scholar.google.com/scholar?q=Learning+and+Generalization+of+Motor+Skills+by+Learning+from+Demonstration
13. Generative Adversarial Imitation Learning — J. Ho, S. Ermon, 2016
https://scholar.google.com/scholar?q=Generative+Adversarial+Imitation+Learning
14. Trust Region Policy Optimization — J. Schulman, S. Levine, P. Moritz, M. Jordan, P. Abbeel, 2015
https://scholar.google.com/scholar?q=Trust+Region+Policy+Optimization
15. Cooperative Inverse Reinforcement Learning — D. Hadfield-Menell, S. J. Russell, P. Abbeel, A. Dragan, 2016
https://scholar.google.com/scholar?q=Cooperative+Inverse+Reinforcement+Learning
16. Time-Contrastive Networks: Self-Supervised Learning from Video — P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, 2017
https://scholar.google.com/scholar?q=Time-Contrastive+Networks%3A+Self-Supervised+Learning+from+Video
17. Deep Q-learning from Demonstrations / Sun et al. on sample-complexity of imitation vs RL — W. Sun, A. Venkatraman, G. Gordon, B. Boots, J. A. Bagnell, 2017
https://scholar.google.com/scholar?q=Deep+Q-learning+from+Demonstrations+%2F+Sun+et+al.+on+sample-complexity+of+imitation+vs+RL
Interactive Visualization: An Algorithmic Perspective on Imitation Learning