1 00:00:01,000 --> 00:00:53,244 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into MPC-Net: A First Principles Guided Policy Search, by Jan Carius et al. — three authors total, Jan Carius, Farbod Farshidian, and Marco Hutter — out of the Robotic Systems Lab at ETH Zürich, submitted to arXiv in September 2019 and accepted at IEEE Robotics and Automation Letters in January 2020. Here's the number that got me: once trained, their policy evaluates in about 0.125 milliseconds on a quadruped robot's onboard computer. The model predictive controller it's replacing takes 38 milliseconds per update. That's a three-hundred-times speedup, and they get there from less than ten minutes of demonstration data. 2 00:00:53,244 --> 00:01:29,142 [Dr. Ada Shannon] And that ten-minute number is the real headline, Hal. Most learning-based legged locomotion work runs on millions of simulated timesteps. This paper is asking a narrower, sharper question: instead of training a policy to copy what an expert controller outputs, can you train it to directly minimize the same mathematical quantity — the control Hamiltonian — that the expert controller is minimizing internally? If that works, you're not imitating behavior, you're imitating the reasoning. That's a meaningfully different bet than standard behavioral cloning, and it's worth understanding why they thought it was necessary in the first place. 3 00:01:29,142 --> 00:01:40,567 [Hal Turing] Okay, so let's back up, because I want to make sure everyone's tracking. Model Predictive Control — MPC — what is it, and why can't you just run it directly on the robot instead of learning a policy at all? 4 00:01:40,567 --> 00:02:43,261 [Dr. Ada Shannon] MPC is an online optimal-control method. At every control tick, it solves a finite-horizon trajectory optimization problem from the robot's current state — given the dynamics model, find the sequence of joint torques or forces that minimizes some cost, like tracking a reference gait, subject to physics constraints. It's powerful because it has an explicit model of the system and can reason about constraints directly. The problem is cost: solving that optimization is expensive, and on a legged robot like ANYmal — ETH's own quadruped platform — you need updates fast enough to react to a foot slipping or a gust of wind. Hwangbo et al., also out of Hutter's lab, published in Science Robotics in 2019, showed reinforcement learning trained entirely in simulation could match or beat hand-tuned MPC controllers once deployed — but that took massive simulated rollout counts to get there. MPC-Net is trying to get MPC's per-step optimality without paying MPC's per-step compute bill, and without RL's appetite for data. 5 00:02:43,261 --> 00:02:59,747 [Hal Turing] Right, and that's where imitation learning comes in — you're using the MPC solver as a teacher and training a network to mimic it. But wait, what's actually being minimized here? Because I keep seeing 'Hamiltonian' in the abstract and I don't have great intuition for what that means outside of physics class. 6 00:02:59,747 --> 00:03:58,075 [Dr. Ada Shannon] Here's the consequence first: minimizing the Hamiltonian instead of copying the expert's raw output gives you a policy that's provably closer to the true optimal control, not just closer to what the teacher happened to say. Now the context — the control Hamiltonian is a quantity from optimal control theory, it comes out of the Hamilton-Jacobi-Bellman equation, and it bundles together the running cost, the constraint penalties, and how a candidate action would push the system's value function forward in time. At the true optimum, the Hamiltonian is at its minimum with respect to the control input. So instead of training the network to match the expert's numbers — a classic supervised regression setup, basically Behavioral Cloning, where expert state-action pairs are treated like labeled training data — you train it to output whatever action drives that Hamiltonian down. It's a subtle but real distinction: matching outputs versus matching the optimality condition those outputs came from— 7 00:03:58,075 --> 00:04:09,407 [Hal Turing] Oh wait wait wait — sorry to cut you off, but does that mean their network never actually sees the expert's chosen action during training? Like, it's not shown 'here's what MPC did, copy it' at all? 8 00:04:09,407 --> 00:05:03,323 [Dr. Ada Shannon] Correct, and that's genuinely the elegant part. The learner is never handed u-star, the optimal control input, directly. It only sees the ingredients needed to evaluate the Hamiltonian at a given state — the value function's derivative, the constraint multipliers — and it's told to minimize that quantity itself. Which connects to something you'll want to compare this against: classical Guided Policy Search, from Levine and Koltun, ICML 2013 — the paper this title is explicitly nodding to with 'a first principles guided policy search.' In that line of work, the trajectory-optimization teacher adapts toward the student as training progresses, so the demonstrations stay within reach of what the network can currently produce. MPC-Net does the opposite. The MPC teacher here never adapts to the learner at all — it just keeps solving the same optimal control problem regardless of how good the student is. 9 00:05:03,323 --> 00:05:24,268 [Hal Turing] Hold on, isn't that a disadvantage though? If the teacher never adjusts to the student's current skill level, doesn't that mean early in training the student is getting demonstrations way outside what it can actually use yet? An adaptive teacher, like Levine and Koltun's setup, seems like it would converge faster because it's meeting the student where it is. 10 00:05:24,268 --> 00:06:02,302 [Dr. Ada Shannon] I actually disagree with you there, Hal. Think about what adaptation costs you: every time the teacher's cost function gets a penalty term added to stay close to the student, you're solving a different, moving-target optimization problem at every iteration, and your demonstrations are only valid for that one snapshot of the student. Here, because the MPC teacher never changes, every single trajectory it ever generates stays valid forever and can be reused indefinitely in the replay buffer. That's a sampling-efficiency argument, not a convergence-speed argument — and given they're claiming under ten minutes of real demonstration data, reuse is clearly what they're optimizing for. 11 00:06:02,302 --> 00:06:23,525 [Hal Turing] Okay, that's fair — I was thinking about it purely from a 'does the student get achievable targets' angle, and you're thinking about it from a 'how many times can I squeeze value out of one expensive MPC solve' angle. I can see both of those mattering, honestly, and I don't think we have to resolve which one wins in the abstract — it probably depends on how expensive your teacher is to query in the first place. 12 00:06:23,525 --> 00:07:05,182 [Dr. Ada Shannon] Agreed, and that's exactly the design constraint they're under — MPC is the expensive part here, so reuse wins. Now, one more piece before we get into the algorithm mechanics: this paper commits to a mixture-of-experts network for the policy itself, not a plain multilayer perceptron. A mixture-of-experts architecture, going back to Jacobs, Jordan, Nowlan, and Hinton's 1991 paper on adaptive mixtures of local experts, has several small sub-networks — 'experts' — each proposing an action, and a gating network that blends their outputs, letting different experts specialize on different situations rather than forcing one network to average across all of them. 13 00:07:05,182 --> 00:07:26,869 [Hal Turing] And the reason that matters specifically for a legged robot is the multimodality, right? Like, a quadruped trotting has completely different contact patterns than one doing a static walk — different legs in the air at different times — so you'd almost want a different specialist for each mode instead of one net trying to smoothly interpolate between totally different behaviors. 14 00:07:26,869 --> 00:08:04,810 [Dr. Ada Shannon] Exactly, and optimal control problems like this one are inherently non-unique — there can be multiple equally valid solutions for the same state, the same way stepping around an obstacle left or right can both be optimal. A single monolithic network forced to average those solutions can produce something dynamically nonsensical. ANYmal itself, Hutter's lab's 18-degree-of-freedom quadruped, is exactly the kind of hybrid, contact-switching system where that multimodality shows up constantly. That's the setup — next, let's get into how the actual training loop pulls demonstrations from MPC and turns them into gradient updates. 15 00:08:04,810 --> 00:08:23,433 [Hal Turing] And the reason that matters specifically for a legged robot is the multimodality, right? Like, a quadruped trotting has completely different contact patterns than one doing a static walk — different legs in the air at different times — so you'd almost want a different specialist for each mode instead of one network trying to average them. 16 00:08:23,433 --> 00:09:09,919 [Dr. Ada Shannon] Exactly, and optimal control problems like this one are inherently non-unique — there can be multiple equally valid solutions for the same state, the same way stepping around an obstacle left or right can both be optimal. A single monolithic network forced to average those solutions can produce something genuinely bad, a policy stuck straddling both options. So each expert gets to own a mode instead. And that feeds straight into how the training loop works: MPC keeps solving fresh trajectory optimizations from random start states, every sample it produces along the way gets dropped into a replay buffer, and the policy — running on a totally separate rhythm — draws random batches from that buffer and takes gradient steps pushing each expert to minimize the Hamiltonian at those points. 17 00:09:09,919 --> 00:09:20,043 [Hal Turing] So MPC and policy training aren't locked in lockstep — MPC refreshes the buffer only occasionally, and the gradient updates just keep grinding on whatever's already sitting there? 18 00:09:20,043 --> 00:10:02,071 [Dr. Ada Shannon] Right. And there's real theory backing why minimizing the Hamiltonian is the correct target rather than just a convenient one — Lemma 1 in the paper bounds how far your policy's output can be from the true optimal control, in terms of how large the Hamiltonian's optimality gap still is at that point. I won't walk through the proof, but the punchline is: drive the Hamiltonian down and you're provably driving the policy toward u-star. And they don't only sample on the nominal trajectory either — they sample in a tube around it, drawn from a Gaussian centered on the nominal state, since SLQ's local value-function approximation stays reasonably accurate nearby. That turns each MPC solve into much more usable data than a single point. 19 00:10:02,071 --> 00:10:15,910 [Hal Turing] Okay, but if MPC is generating all this from its own optimal rollouts, and the policy's never tested on the states its own mistakes would actually put it in, doesn't the training distribution just drift away from reality once the policy's running on its own? 20 00:10:15,910 --> 00:10:42,102 [Dr. Ada Shannon] That's exactly the distribution mismatch problem, and their fix is a direct adaptation of Ross, Gordon, and Bagnell's DAgger work, Carnegie Mellon, 2011 — a no-regret reduction of imitation learning to online learning. They're upfront it's borrowed machinery, not a new contribution. Equations 12 and 13 define a behavioral policy as a weighted blend of MPC and whatever the current network out— 21 00:10:42,102 --> 00:10:53,759 [Hal Turing] Wait, wait — sorry, cutting you off — so the thing actually steering the simulated rollout during data collection isn't pure MPC the whole time? It's partly the network's own output, even mid-training? 22 00:10:53,759 --> 00:11:27,428 [Dr. Ada Shannon] Right, that's the trick. Early on, the mixing parameter alpha is near zero, so MPC does almost all the steering. By the final iteration alpha equals one, and the learned policy fully decides where the rollout goes — though MPC is still the one asked what the optimal action is at whatever state that turns out to be. So you get MPC's Hamiltonian labels, sampled from the distribution the imperfect student would actually wander into. Worth noting: MPC itself is never influenced by the student, only the states it gets evaluated at shift. 23 00:11:27,428 --> 00:11:40,292 [Hal Turing] So let's get into the mixture-of-experts piece itself then — the output's a convex combination of expert sub-policies. How does the gating network actually decide the weights, and why not just use softmax? 24 00:11:40,292 --> 00:12:23,156 [Dr. Ada Shannon] Softmax technically satisfies the constraint that weights are nonnegative and sum to one, but they found a sigmoid activation followed by normalization works better in practice, because softmax is too decisive — it sharply commits to one expert per state, and an unlucky initialization can mean some expert never gets picked early on and simply dies, no gradient ever reaches it again. Sigmoid-then-normalize spreads responsibility more evenly, so training consistently ends up using a sensible number of experts instead of losing some to bad luck. And critically, they don't just train the combined output to minimize the Hamiltonian — each individual expert is forced to minimize it too, which is what drives real specialization instead of redundant copies. 25 00:12:23,156 --> 00:12:29,425 [Hal Turing] Give me the actual numbers, Ada — what does this look like in practice, training length, model size, that kind of thing? 26 00:12:29,425 --> 00:13:13,543 [Dr. Ada Shannon] A hundred thousand iterations total, batch size 32, learning rate 1e-3, replay buffer holding 100,000 samples, MPC refreshing it every 500 iterations. The kinodynamic model is 24 states — base pose and twist plus joint angles — and 24 inputs, joint velocities and contact forces. Then Table II runs a narrower sanity check: forty points on or near an optimal trajectory, comparing MPC's actual output against just minimizing the Hamiltonian directly. Constraint violation for pure MPC is essentially zero; the Hamiltonian minimization comes in a bit higher but still tiny, with relative deviation from the true optimal control under half a percent. 27 00:13:13,543 --> 00:13:23,481 [Hal Turing] That lines up with the behavioral cloning comparison too — Figure 4 shows the Hamiltonian loss beating a plain BC loss that just matches MPC's raw output? 28 00:13:23,481 --> 00:13:56,036 [Dr. Ada Shannon] Right, and the reported failure mode is almost darkly funny — the BC policy, dropped into the simulator, tends to fall over after a few footsteps because it's been quietly accumulating constraint violations the whole time it looked stable. Their read is that it was essentially cheating, faking stability by ignoring physical constraints, until the cheating became unsustainable. One caveat though: that specific claim is qualitative — no trial count, no success rate, no variance, just an anecdote sitting next to otherwise quantitative figures. 29 00:13:56,036 --> 00:14:02,398 [Hal Turing] And the efficiency number from the abstract — under ten minutes of demonstration data — where does that actually come from? 30 00:14:02,398 --> 00:14:32,955 [Dr. Ada Shannon] The tube sampling is what gets them there — policies become usable on the robot at around seventy-five percent of total training iterations, meaning the buffer backing that usable policy is equivalent to about nine minutes of an optimal controller actually running the robot. And once trained, evaluating the policy onboard takes roughly 0.125 milliseconds versus 38 milliseconds for a full MPC update — about three hundred times faster, which is what lets it run synchronously with the tracking controller. 31 00:14:32,955 --> 00:14:49,116 [Hal Turing] Here's the part I find almost too clean, though — Figure 6 shows the gating network switching experts exactly when contact configuration changes, three experts for trotting, four for static walk. That's the network discovering gait structure entirely on its own. 32 00:14:49,116 --> 00:15:09,085 [Dr. Ada Shannon] I actually disagree with you there, Hal. The input isn't raw time — it's four phase variables, one per leg, already zeroed during stance and tracing a sine wave through swing. That's handing the network an explicit contact-phase signal. Of course the gating aligns with contact changes; the input is basically screaming it. 33 00:15:09,085 --> 00:15:18,420 [Hal Turing] Sure, but knowing the phase doesn't force the gating to carve experts along contact boundaries specifically — it could've split some other arbitrary way and still had the same inputs available. 34 00:15:18,420 --> 00:15:33,699 [Dr. Ada Shannon] Fair, I'll give you that much — landing on the physically sensible partition rather than an arbitrary one is real credit. I just wouldn't call it discovery from scratch when the phase signal was doing a lot of the heavy lifting already. Strongly hinted discovery, not emergent. 35 00:15:33,699 --> 00:15:40,665 [Hal Turing] I can live with strongly hinted. So does any of this actually carry over onto the real hardware, or is that a separate story entirely? 36 00:15:40,665 --> 00:16:06,253 [Dr. Ada Shannon] They test both gaits directly on the physical ANYmal, same network structure and hyperparameters for each. They displace the robot from its nominal position and yaw, then let the policy drive it back. The result is qualitative but clean — base position and yaw both return to zero with minimal overshoot, and the onboard policy runs fast enough to be called synchronously with the tracking controller the entire time. 37 00:16:06,253 --> 00:16:30,959 [Dr. Ada Shannon] Overshoot, yeah — and it's compelling, but I want to be honest about what kind of evidence this actually is. Two gaits, on flat ground, no trial count, no disturbance magnitude reported, no comparison against MPC running on the same hardware under the same push. It's basically 'here's a plot, the line goes back to zero.' That's the softest evidence in the entire paper, sitting right next to the tightest theory in Section two. 38 00:16:30,959 --> 00:16:59,055 [Hal Turing] Which is exactly the tension I want to dig into, Ada. Go back to Table II — the Hamiltonian-versus-optimal-control comparison. That's forty points. From one system. And the text around it basically claims minimizing the Hamiltonian 'yields optimal controls,' phrased like a general finding. Forty points on one quadruped's kinodynamic model reads like a sanity check to me, not validation of a general claim. 39 00:16:59,055 --> 00:17:34,860 [Dr. Ada Shannon] It is a sanity check, and I think that's the honest label. The numbers themselves are reassuring — relative error to u-star around 1.6e-3 on-trajectory, 2.8e-2 near it — small enough that Lemma 1's assumptions look satisfied in practice. But the scope is exactly one 24-state, 24-input model solved with one SLQ solver, the same one from Farshidian, Neunert, Winkler, Rey and Buchli's efficient optimal planning framework paper out of ETH, 2017. Nothing here tells you it holds for a different robot or a different solver family. 40 00:17:34,860 --> 00:17:54,040 [Hal Turing] Wait, hold on — that SLQ dependency is actually bigger than it sounds to me. This whole method leans on the solver handing you the value function derivative and the Lagrange multipliers as a free byproduct. What happens if your MPC is sampling-based, like an MPPI-style controller, which doesn't naturally give you any of that? 41 00:17:54,040 --> 00:18:39,133 [Dr. Ada Shannon] Then you're stuck, honestly. The Hamiltonian loss needs those second-order value terms and multipliers at training time, and that's a DDP-family feature, not a universal MPC feature. So the paper frames itself as general for 'dynamical systems with a known model,' but functionally it's general for systems solved with an SLQ/DDP-style solver, on one robot, two gaits. Compare that to the sim-to-real line — Hwangbo, Lee, Dosovitskiy, Bellicoso, Tsounis, Koltun and Hutter's Science Robotics paper, 2019, or Tan et al.'s sim-to-real quadruped work, 2018 — those get tested across varied terrain and disturbance regimes. MPC-Net gets two gaits on presumably benign ground. 42 00:18:39,133 --> 00:19:06,626 [Hal Turing] Okay, but here's where I actually push back a little. The paper says outright, in its own conclusion, that by design this method cannot outperform MPC because it's optimizing the exact same cost function. If your ceiling is just 'as good as the MPC solution, including whatever local minimum MPC got stuck in,' doesn't that undercut the entire value proposition? You've built an elegant loss function to reproduce a controller you already have. 43 00:19:06,626 --> 00:19:30,774 [Dr. Ada Shannon] I actually disagree with you there, Hal. The value was never 'produce a better controller than MPC' — it was always speed. The whole point is you get that same solution quality at 0.125 milliseconds instead of 38, without the online solve. That's not a footnote, that's the entire product. A ceiling equal to MPC is fine if MPC's ceiling is already good enough for the task and the bottleneck was always compute. 44 00:19:30,774 --> 00:19:48,932 [Hal Turing] Sure, but 'good enough for the task' is doing a lot of work in that sentence. If MPC's local minimum is bad on a given terrain, you've now baked that mediocrity into a policy that runs forever and has no way to escape it — no exploration mechanism, by the paper's own admission. 45 00:19:48,932 --> 00:20:29,985 [Dr. Ada Shannon] That's fair, and I'll meet you there — it's a real limitation, not a strawman. Where we land, I think, is that this is an honest tradeoff rather than hype: you're trading ceiling for latency, and the paper says so plainly instead of hiding it. Same honesty question applies to the mixture-of-experts claim in Section three-E — 'significantly better constraint violation' than an MLP, but that's five averaged runs and a figure you eyeball, not a statistical test. The architecture itself is a direct application of Jacobs, Jordan, Nowlan and Hinton's Adaptive Mixtures of Local Experts, 1991 — old idea, competently reused, not reinvented. 46 00:20:29,985 --> 00:20:59,242 [Hal Turing] Which loops back to something worth naming plainly: the genuinely new piece here isn't the teacher-student setup — that's Levine and Koltun's Guided Policy Search lineage — and it isn't the distribution-matching fix, which borrows straight from Ross, Gordon and Bagnell's DAgger. The actual contribution is the loss function itself, minimizing the Hamiltonian instead of a distance metric. Everything else is solid engineering built on existing pieces. 47 00:20:59,242 --> 00:21:36,162 [Dr. Ada Shannon] Right, and worth noting — Marco Hutter, one of the three authors here, later co-authored 'Robots Need More than VLA and World Models,' arguing embodied systems need more than end-to-end learned models. You can see that skepticism about black-box policies germinating here: MPC-Net keeps a model-based optimal-control core and only learns to compress it. Practically, that 0.125-millisecond number matters most for teams hitting a control-frequency wall MPC can't clear, or wanting to learn directly on hardware instead of in sim. It's not a drop-in replacement for RL pipelines chasing robustness across terrain. 48 00:21:36,162 --> 00:22:05,234 [Hal Turing] So, wrapping up — Carius, Farshidian and Hutter give you a genuinely elegant piece of theory in the Hamiltonian-minimization loss, backed by a real bound in Lemma 1, but the empirical case is thin: one robot, one solver family, two gaits, small sample counts throughout, and a hardware validation that's really just 'it didn't fall.' Good idea, undersold evidence. Thanks for listening, everyone — we'll catch you next time.