JEPA + Model Predictive Control
A joint-embedding predictor forecasts the next latent state instead of raw pixels; MPC uses that predictor to plan, but only ever executes the first action of the plan.

Prediction pipeline & the adaptation signal

The encoder never reconstructs pixels — it only has to preserve what the predictor needs. The prediction loss between the predicted and actually-encoded latent is the only signal used both for pretraining and for the live test-time update.

Receding-horizon MPC, step by step

Replanning step 0

At every step the model imagines several candidate rollouts (purple), picks the lowest-cost one (green), executes only its first action, then throws the rest away and replans from the new state — the same receding-horizon idea used in refineries since the 1980s and to land Falcon 9 boosters.

Plan → Execute → Observe → Adapt
Test-time adaptation threads directly into the control loop: every observed transition becomes a free, label-free training example.

The closed loop

No reward signal, no human in the loop, no separate data-collection phase — the loop runs every single step of the episode, borrowing the recipe from image-classification test-time training (Sun et al., 2020) and echoing cerebellar motor recalibration.

Replay buffer: recent-5 vs hard-N

Target latents are stop-gradient — the model cannot cheat by dragging both sides toward each other. Buffer selection policy barely moves the needle; what matters is that adaptation happens at all.

Which layers get updated

Default is predlast+enclast — just the predictor’s last transformer block plus the encoder’s last stage, everything else frozen. LoRA-everywhere, first-layer-only, and encoder-first-stage all beat the frozen baseline too — the win is the adaptation itself, not a magic subset of parameters.

Four kinds of distribution shift
Geometry, perception, physics, and topology — tested separately, on PushObj / PushT and PointMaze-Medium.

Frozen vs. adapted, by shift type

Unseen object shapes basically double in success rate under adaptation. Color recoloring is the soft spot: the frozen model leaned on color as a shortcut, and a couple of adapted layers cannot undo that.

Condition-level breakdown

Gain (rightmost column) runs 19–32 points everywhere except the three color-recolor conditions, which stay stuck around 7–8 points — the one visible dark cell in an otherwise warm grid.

Does it actually deliver?
Success climbs across replanning steps, generalizes across backbones, and the biggest win shows up exactly where training data is scarce.

Success rate across replanning steps

The frozen model plateaus almost immediately; the adapted model keeps climbing for the full 30-step episode.

Data efficiency

A model trained on one shape with 1,000 trajectories, adapted at test time, beats a frozen model trained on four shapes with 16,000 trajectories each — 16× more data and it still loses.

Generalizes across backbones

Every backbone tested — AdaJEPA, DINO-WM, Temporal Straightening — improves under adaptation.

Added latency per replan

One to three hundredths of a second on an H200 — the adaptation step is essentially free next to a control loop that is already replanning constantly.

What the paper does not show
Every headline comparison is AdaJEPA versus its own frozen checkpoint. Four cited prior methods never appear in a results table.

Uncompared prior work

AdaWM (driving), AdaWorld, Lanier et al.’s latent-state dynamics residuals, and Parthasarathy et al.’s train-test gap paper are all named as related work — none are run as a baseline. What is demonstrated is one gradient step beats a frozen checkpoint, not that it beats the existing toolkit.

Stability is a narrow peak, not a safe default

Push the learning-rate multiplier up or down from the shipped default and success drops — fine for one throwaway episode, a bad sign for accumulating updates across a robot’s whole lifetime.

Color invariance stays bounded, every variant

Shape-shift success climbs sharply under any adapted variant; color-shift success barely moves under any of them — a frozen encoder that never learned color invariance cannot have it patched in by a light last-layer nudge.

References

  1. 1.AdaJEPA: An Adaptive Latent World Model — Ying Wang, Oumayma Bounou, Yann LeCun, Mengye Ren, 2026
    arxiv.org/abs/2606.32026
  2. 2.Model Predictive Control: Theory and Practice — A Survey — Carlos E. García, David M. Prett, Manfred Morari, 1989
    Google Scholar
  3. 3.Model Predictive Control: Classical, Robust and Stochastic — Basil Kouvaritakis, Mark Cannon, 2016
    Google Scholar
  4. 4.Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models (PETS) — Kurtland Chua, Roberto Calandra, Rowan McAllister, Sergey Levine, 2018
    Google Scholar
  5. 5.Learning Latent Dynamics for Planning from Pixels (PlaNet) — Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, James Davidson, 2019
    Google Scholar
  6. 6.Dino-wm: World models on pre-trained visual features enable zero-shot planning — Zhou, G., Pan, H., LeCun, Y., Pinto, L., 2025
    Google Scholar
  7. 7.Temporal straightening for latent planning — Wang, Y., Bounou, O., Zhou, G., Balestriero, R., Rudner, T. G., LeCun, Y., Ren, M., 2026
    Google Scholar
  8. 8.Closing the train-test gap in world models for gradient-based planning — Parthasarathy, A., Kalra, N., Agrawal, R., LeCun, Y., Bounou, O., Izmailov, P., Goldblum, M., 2025
    Google Scholar
  9. 9.Td-mpc2: Scalable, robust world models for continuous control — Hansen, N., Su, H., Wang, X., 2024
    Google Scholar
  10. 10.Adawm: Adaptive world model based planning for autonomous driving — Wang, H., Ye, X., Tao, F., Pan, C., Mallik, A., Yaman, B., Ren, L., Zhang, J., 2025
    Google Scholar
  11. 11.Test-time training with self-supervision for generalization under distribution shifts — Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A. A., Hardt, M., 2020
    Google Scholar