MPC re-solves a constrained trajectory optimization at every control tick. MPC-Net trains a fast network offline against the MPC solver's own optimality condition, then runs that network onboard instead of the solver.
MPC never adapts to the student — it keeps re-solving the same optimal control problem. DAgger-style mixing only changes which states get sampled, not what MPC is optimizing.
Regress directly onto the expert's chosen action u*. Learns surface behavior; constraint violation can accumulate silently.
Never shown u* — only the ingredients (∂V/∂x, constraint multipliers λ) needed to evaluate H(x,u), then trained to drive it down.
Relative deviation from the true optimal control u*, and constraint violation, comparing MPC's actual output against pure Hamiltonian minimization.
Small-sample sanity check, not general validation: one 24-state / 24-input kinodynamic model, solved with a single SLQ solver.
Each leg contributes a phase variable (zero during stance, sinusoidal during swing). The gating network's weights track contact-configuration changes — strongly hinted by that phase input, not free-form emergent discovery.
Hover a line to isolate one leg. Phase is pinned to zero through stance and sweeps through swing — this is the signal the gating network is conditioned on.
α starts near 0 (MPC steers rollouts) and reaches 1 by the final iteration (the learned policy steers rollouts). MPC always supplies the Hamiltonian label — only the sampling distribution shifts.
Base position/yaw error after an external push, running fully onboard at 0.125 ms/step. No trial count or disturbance magnitude reported in the paper — treat this curve as illustrative, not measured.
The theory (Lemma 1) is tight; the empirical support thins out fast toward hardware.