AI Post Transformers · Episode Companion Viz

Next-Latent Prediction Lets Transformers Learn Compact World Models

arXiv:2511.05963 Microsoft Research · 2025 Belief States · Self-Speculative Decoding

A small auxiliary loss — no architecture change — pushes an ordinary transformer to compress its history into a compact belief state. This page visualizes the method's mechanics, the theorem behind the guarantee, the Manhattan / Countdown / Path-Star results, and the self-speculative decoding speedups, alongside where the paper's evidence is still proxy rather than direct.

Four ways to get a compact state architecture comparison

GPT keeps a growing KV cache with no pressure to compress. BST pays for a guarantee with a second encoder. JTP is cheap but its guarantee depends on an unknown horizon k. NextLat adds a small dynamics network on the side.

Effective rank on Manhattan lower = more compact

Exponentiated entropy of the hidden states' singular value spectrum. All models hit 100% next-turn accuracy — this is what actually separates them.

NextLat 52.7 GPT 160.1 JTP 215.8

Manhattan taxi metrics

Valid out-of-distribution routes, sequence compression (do two paths to the same intersection continue identically?), and robustness when a street is closed (detour).

The Clever Hans problem this is fixing

A transformer can reach 100% next-turn accuracy while its reconstructed internal map is incoherent — it drives fine until a street closes, because next-token accuracy never tested whether a real map existed.

Backward induction: why ht becomes a belief state

If decoding from ht matches the true next-token law (next-token consistency), and the dynamics model matches the true law of the transformer's own next hidden state (transition consistency), then decode→update→repeat needs no prefix — the last step anchors the induction, longer horizons only add signal.

Three loss terms, one guard against collapse

Predicting your own future state trivially collapses to a constant unless something stops it. Stop-gradient on the target + cross-entropy on tokens are the two guards.

Two conditions the theorem needs

Neither condition is measured directly in the paper — Appendix E shows the Smooth-L1 latent loss actually rising during learning-rate cooldown under both AdamW and Muon.

Open gap: nobody probes sufficiency directly — feed ht alone (no attention context) and check whether next-k predictions survive.

Countdown accuracy by rollout horizon d

BST and GPT numbers only exist at horizon 1 — BST doesn't scale to Manhattan or the 1.3B run, so the strongest baseline mostly appears where it loses.

Perplexity & zero-shot accuracy 1.3B / 100B tokens FineWeb-Edu

Perplexity still favors GPT. The zero-shot accuracy edge is a single seed across 9 tasks — the authors call it inconsistent.

Path-Star accuracy heatmap

Hover a cell. G(7,7) is the harder, original Bachmann & Nagarajan setup — not comparable to the BST/JTP papers' own numbers.

BST's G(7,7) failure rate is not reported precisely in the paper; the low value shown is illustrative of "fails," not a quoted figure.

A5 group-multiplication control NC1-complete, out of transformer reach

A 2-layer transformer trained at 12 tokens fails at 36. A small MLP, co-trained by regressing onto the transformer's states, reaches 95% at 36 tokens — with far fewer parameters than the transformer it learned from.

Self-speculative decoding via the dynamics model

The MLP recurses on its own predicted states to draft tokens of any length, then the transformer verifies the whole draft in one parallel pass — unlike MTP heads, which freeze draft length at the training horizon.

Decoding speedup by domain horizon = 2

JTP and MTP figures are single averages reported in the paper, not measured per-domain — shown here as flat reference lines against NextLat's per-domain bars.

Training cost vs. the guarantee it buys

References

  1. Next-Latent Prediction Transformers Learn Compact World Models — Teoh et al., 2025. arXiv:2511.05963
  2. Bootstrap Your Own Latent — Grill et al., 2020. scholar
  3. Data-Efficient RL with Self-Predictive Representations — Schwarzer et al., 2021. scholar
  4. Better & Faster LLMs via Multi-token Prediction — Gloeckle et al., 2024. scholar
  5. A Path Towards Autonomous Machine Intelligence (JEPA) — LeCun, 2022. scholar
  6. Planning and Acting in Partially Observable Stochastic Domains — Kaelbling, Littman, Cassandra, 1998. scholar
  7. Predictive Representations of State — Littman, Sutton, Singh, 2001. scholar
  8. The Belief State Transformer — Hu et al., 2024. scholar
  9. Transformers Represent Belief State Geometry in their Residual Stream — Shai et al., 2024. scholar
  10. Fast Inference from Transformers via Speculative Decoding — Leviathan, Kalman, Matias, 2023. scholar
  11. Accelerating LLM Decoding with Speculative Sampling — Chen et al., 2023. scholar
  12. Medusa: Multiple Decoding Heads — Cai et al., 2024. scholar
  13. EAGLE: Rethinking Feature Uncertainty — Li et al., 2024. scholar
  14. Efficient Joint Prediction of Multiple Future Tokens (JTP) — Ahn, Lamb, Langford, 2025. scholar
  15. Bridging State and History Representations — Ni et al., 2024. scholar
  16. The Pitfalls of Next-Token Prediction — Bachmann, Nagarajan, 2024. scholar
  17. Evaluating the World Model Implicit in a Generative Model — Vafa et al., 2024. scholar
  18. EAGLE-2 — Li et al., 2024. scholar
  19. Self-Predictive Representations (SPR) — Schwarzer et al., 2021. scholar
  20. The Parallelism Tradeoff — Merrill, Sabharwal, 2023. scholar
  21. The Illusion of State in State-Space Models — Merrill, Petty, Sabharwal, 2025. scholar
  22. LLM-JEPA — Huang, LeCun, Balestriero, 2025. scholar
  23. Belief State Geometry in the Residual Stream — Shai et al., 2024. scholar
  24. DeepSeek-V3 Technical Report (MTP) — DeepSeek-AI, 2024. scholar

Interactive visualization companion: full viz page