This episode examines MPC-Net, a 2019/2020 paper from ETH Zürich's Robotic Systems Lab that trains a fast neural policy to replace expensive model predictive control on the ANYmal quadruped, cutting per-step evaluation from 38 milliseconds to roughly 0.125 milliseconds using less than ten minutes of demonstration data. The discussion centers on why the method learns by minimizing the control Hamiltonian — the optimality condition MPC itself solves internally — rather than copying the expert's chosen actions, arguing this teaches the network the underlying reasoning rather than surface behavior. It contrasts this approach with classical Guided Policy Search, where the teacher adapts toward the student over training, versus MPC-Net's fixed, non-adaptive teacher that keeps solving the same optimal control problem regardless of the learner's progress. The hosts debate the tradeoffs of adaptive versus static teachers in imitation learning, weighing convergence speed against the validity and reusability of generated trajectories. Listeners interested in legged robotics, optimal control theory, or the mechanics of imitation learning will find a detailed technical walkthrough of how theory-grounded objectives can outperform standard behavioral cloning.
Sources:
1. MPC-Net: A First Principles Guided Policy Search — Jan Carius, Farbod Farshidian, Marco Hutter, 2019
http://arxiv.org/abs/1909.051972. ALVINN: An Autonomous Land Vehicle in a Neural Network — Dean Pomerleau, 1989
https://scholar.google.com/scholar?q=ALVINN%3A+An+Autonomous+Land+Vehicle+in+a+Neural+Network3. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%294. Guided Policy Search — Sergey Levine, Vladlen Koltun, 2013
https://scholar.google.com/scholar?q=Guided+Policy+Search5. Generative Adversarial Imitation Learning — Jonathan Ho, Stefano Ermon, 2016
https://scholar.google.com/scholar?q=Generative+Adversarial+Imitation+Learning6. Learning Agile and Dynamic Motor Skills for Legged Robots — Jemin Hwangbo, Joonho Lee, Alexey Dosovitskiy, Dario Bellicoso, Vassilios Tsounis, Vladlen Koltun, Marco Hutter, 2019
https://scholar.google.com/scholar?q=Learning+Agile+and+Dynamic+Motor+Skills+for+Legged+Robots7. Sim-to-Real: Learning Agile Locomotion for Quadruped Robots — Jie Tan, Tingnan Zhang, Erwin Coumans, Atil Iscen, Yunfei Bai, Danijar Hafner, Steven Bohez, Vincent Vanhoucke, 2018
https://scholar.google.com/scholar?q=Sim-to-Real%3A+Learning+Agile+Locomotion+for+Quadruped+Robots8. Learning Quadrupedal Locomotion over Challenging Terrain — Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen, Vladlen Koltun, Marco Hutter, 2020
https://scholar.google.com/scholar?q=Learning+Quadrupedal+Locomotion+over+Challenging+Terrain9. High-Slope Terrain Locomotion for Torque-Controlled Quadruped Robots (representative MPC/whole-body baseline) — Marco Hutter, Christian Gehring, and colleagues, ETH Zurich Robotic Systems Lab, 2016-2018 (various)
https://scholar.google.com/scholar?q=High-Slope+Terrain+Locomotion+for+Torque-Controlled+Quadruped+Robots+%28representative+MPC%2Fwhole-body+baseline%2910. Adaptive mixtures of local experts — R. A. Jacobs, M. I. Jordan, S. J. Nowlan, G. E. Hinton, 1991
https://scholar.google.com/scholar?q=Adaptive+mixtures+of+local+experts11. An efficient optimal planning and control framework for quadrupedal locomotion — F. Farshidian, M. Neunert, A. W. Winkler, G. Rey, J. Buchli, 2017
https://scholar.google.com/scholar?q=An+efficient+optimal+planning+and+control+framework+for+quadrupedal+locomotionInteractive Visualization: MPC-Net: Learning Optimal Control via the Hamiltonian