MPC-Net: Learning Optimal Control via the Hamiltonian

Jan Carius, Farbod Farshidian, Marco Hutter — Robotic Systems Lab, ETH Zürich
Submitted arXiv September 2019 · IEEE Robotics and Automation Letters, January 2020
arXiv:1909.05197 AI Post Transformers ANYmal quadruped · SLQ / DDP-based MPC

Why replace MPC at all?

MPC re-solves a constrained trajectory optimization at every control tick. MPC-Net trains a fast network offline against the MPC solver's own optimality condition, then runs that network onboard instead of the solver.

38 ms
MPC solve, per step
0.125 ms
MPC-Net policy, per step
~304×
speedup
<10 min
demonstration data used

Training pipeline

MPC never adapts to the student — it keeps re-solving the same optimal control problem. DAgger-style mixing only changes which states get sampled, not what MPC is optimizing.

data / gradient flow DAgger rollout mixing (α: 0→1)

Matching outputs vs. matching optimality

Behavioral Cloning

Regress directly onto the expert's chosen action u*. Learns surface behavior; constraint violation can accumulate silently.

Hamiltonian Minimization (MPC-Net)

Never shown u* — only the ingredients (∂V/∂x, constraint multipliers λ) needed to evaluate H(x,u), then trained to drive it down.

Table II sanity check — 40 points, one robot, one solver

Relative deviation from the true optimal control u*, and constraint violation, comparing MPC's actual output against pure Hamiltonian minimization.

Small-sample sanity check, not general validation: one 24-state / 24-input kinodynamic model, solved with a single SLQ solver.

Gating weight across the gait cycle

Each leg contributes a phase variable (zero during stance, sinusoidal during swing). The gating network's weights track contact-configuration changes — strongly hinted by that phase input, not free-form emergent discovery.

low gating weight → high gating weight

Per-leg phase signal (input to the gate)

Hover a line to isolate one leg. Phase is pinned to zero through stance and sweeps through swing — this is the signal the gating network is conditioned on.

DAgger mixing schedule

α starts near 0 (MPC steers rollouts) and reaches 1 by the final iteration (the learned policy steers rollouts). MPC always supplies the Hamiltonian label — only the sampling distribution shifts.

Hardware push-recovery (qualitative)

Base position/yaw error after an external push, running fully onboard at 0.125 ms/step. No trial count or disturbance magnitude reported in the paper — treat this curve as illustrative, not measured.

Evidence strength, by claim

The theory (Lemma 1) is tight; the empirical support thins out fast toward hardware.

References