A joint-embedding predictor forecasts the next latent state instead of raw pixels; MPC uses that predictor to plan, but only ever executes the first action of the plan.
Prediction pipeline & the adaptation signal
The encoder never reconstructs pixels — it only has to preserve what the predictor needs. The prediction loss between the predicted and actually-encoded latent is the only signal used both for pretraining and for the live test-time update.
Receding-horizon MPC, step by step
Replanning step 0
At every step the model imagines several candidate rollouts (purple), picks the lowest-cost one (green), executes only its first action, then throws the rest away and replans from the new state — the same receding-horizon idea used in refineries since the 1980s and to land Falcon 9 boosters.
Plan → Execute → Observe → Adapt
Test-time adaptation threads directly into the control loop: every observed transition becomes a free, label-free training example.
The closed loop
No reward signal, no human in the loop, no separate data-collection phase — the loop runs every single step of the episode, borrowing the recipe from image-classification test-time training (Sun et al., 2020) and echoing cerebellar motor recalibration.
Replay buffer: recent-5 vs hard-N
Target latents are stop-gradient — the model cannot cheat by dragging both sides toward each other. Buffer selection policy barely moves the needle; what matters is that adaptation happens at all.
Which layers get updated
Default is predlast+enclast — just the predictor’s last transformer block plus the encoder’s last stage, everything else frozen. LoRA-everywhere, first-layer-only, and encoder-first-stage all beat the frozen baseline too — the win is the adaptation itself, not a magic subset of parameters.
Four kinds of distribution shift
Geometry, perception, physics, and topology — tested separately, on PushObj / PushT and PointMaze-Medium.
Frozen vs. adapted, by shift type
Unseen object shapes basically double in success rate under adaptation. Color recoloring is the soft spot: the frozen model leaned on color as a shortcut, and a couple of adapted layers cannot undo that.
Condition-level breakdown
Gain (rightmost column) runs 19–32 points everywhere except the three color-recolor conditions, which stay stuck around 7–8 points — the one visible dark cell in an otherwise warm grid.
Does it actually deliver?
Success climbs across replanning steps, generalizes across backbones, and the biggest win shows up exactly where training data is scarce.
Success rate across replanning steps
The frozen model plateaus almost immediately; the adapted model keeps climbing for the full 30-step episode.
Data efficiency
A model trained on one shape with 1,000 trajectories, adapted at test time, beats a frozen model trained on four shapes with 16,000 trajectories each — 16× more data and it still loses.
Generalizes across backbones
Every backbone tested — AdaJEPA, DINO-WM, Temporal Straightening — improves under adaptation.
Added latency per replan
One to three hundredths of a second on an H200 — the adaptation step is essentially free next to a control loop that is already replanning constantly.
What the paper does not show
Every headline comparison is AdaJEPA versus its own frozen checkpoint. Four cited prior methods never appear in a results table.
Uncompared prior work
AdaWM (driving), AdaWorld, Lanier et al.’s latent-state dynamics residuals, and Parthasarathy et al.’s train-test gap paper are all named as related work — none are run as a baseline. What is demonstrated is one gradient step beats a frozen checkpoint, not that it beats the existing toolkit.
Stability is a narrow peak, not a safe default
Push the learning-rate multiplier up or down from the shipped default and success drops — fine for one throwaway episode, a bad sign for accumulating updates across a robot’s whole lifetime.
Color invariance stays bounded, every variant
Shape-shift success climbs sharply under any adapted variant; color-shift success barely moves under any of them — a frozen encoder that never learned color invariance cannot have it patched in by a light last-layer nudge.
References
1.AdaJEPA: An Adaptive Latent World Model — Ying Wang, Oumayma Bounou, Yann LeCun, Mengye Ren, 2026 arxiv.org/abs/2606.32026
2.Model Predictive Control: Theory and Practice — A Survey — Carlos E. García, David M. Prett, Manfred Morari, 1989 Google Scholar
3.Model Predictive Control: Classical, Robust and Stochastic — Basil Kouvaritakis, Mark Cannon, 2016 Google Scholar
4.Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models (PETS) — Kurtland Chua, Roberto Calandra, Rowan McAllister, Sergey Levine, 2018 Google Scholar
5.Learning Latent Dynamics for Planning from Pixels (PlaNet) — Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, James Davidson, 2019 Google Scholar
6.Dino-wm: World models on pre-trained visual features enable zero-shot planning — Zhou, G., Pan, H., LeCun, Y., Pinto, L., 2025 Google Scholar
7.Temporal straightening for latent planning — Wang, Y., Bounou, O., Zhou, G., Balestriero, R., Rudner, T. G., LeCun, Y., Ren, M., 2026 Google Scholar
8.Closing the train-test gap in world models for gradient-based planning — Parthasarathy, A., Kalra, N., Agrawal, R., LeCun, Y., Bounou, O., Izmailov, P., Goldblum, M., 2025 Google Scholar
9.Td-mpc2: Scalable, robust world models for continuous control — Hansen, N., Su, H., Wang, X., 2024 Google Scholar
10.Adawm: Adaptive world model based planning for autonomous driving — Wang, H., Ye, X., Tao, F., Pan, C., Mallik, A., Yaman, B., Ren, L., Zhang, J., 2025 Google Scholar
11.Test-time training with self-supervision for generalization under distribution shifts — Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A. A., Hardt, M., 2020 Google Scholar