A small auxiliary loss — no architecture change — pushes an ordinary transformer to compress its history into a compact belief state. This page visualizes the method's mechanics, the theorem behind the guarantee, the Manhattan / Countdown / Path-Star results, and the self-speculative decoding speedups, alongside where the paper's evidence is still proxy rather than direct.
GPT keeps a growing KV cache with no pressure to compress. BST pays for a guarantee with a second encoder. JTP is cheap but its guarantee depends on an unknown horizon k. NextLat adds a small dynamics network on the side.
Exponentiated entropy of the hidden states' singular value spectrum. All models hit 100% next-turn accuracy — this is what actually separates them.
Valid out-of-distribution routes, sequence compression (do two paths to the same intersection continue identically?), and robustness when a street is closed (detour).
A transformer can reach 100% next-turn accuracy while its reconstructed internal map is incoherent — it drives fine until a street closes, because next-token accuracy never tested whether a real map existed.
If decoding from ht matches the true next-token law (next-token consistency), and the dynamics model matches the true law of the transformer's own next hidden state (transition consistency), then decode→update→repeat needs no prefix — the last step anchors the induction, longer horizons only add signal.
Predicting your own future state trivially collapses to a constant unless something stops it. Stop-gradient on the target + cross-entropy on tokens are the two guards.
Neither condition is measured directly in the paper — Appendix E shows the Smooth-L1 latent loss actually rising during learning-rate cooldown under both AdamW and Muon.
Open gap: nobody probes sufficiency directly — feed ht alone (no attention context) and check whether next-k predictions survive.
BST and GPT numbers only exist at horizon 1 — BST doesn't scale to Manhattan or the 1.3B run, so the strongest baseline mostly appears where it loses.
Perplexity still favors GPT. The zero-shot accuracy edge is a single seed across 9 tasks — the authors call it inconsistent.
Hover a cell. G(7,7) is the harder, original Bachmann & Nagarajan setup — not comparable to the BST/JTP papers' own numbers.
BST's G(7,7) failure rate is not reported precisely in the paper; the low value shown is illustrative of "fails," not a quoted figure.
A 2-layer transformer trained at 12 tokens fails at 36. A small MLP, co-trained by regressing onto the transformer's states, reaches 95% at 36 tokens — with far fewer parameters than the transformer it learned from.
The MLP recurses on its own predicted states to draft tokens of any length, then the transformer verifies the whole draft in one parallel pass — unlike MTP heads, which freeze draft length at the training horizon.
JTP and MTP figures are single averages reported in the paper, not measured per-domain — shown here as flat reference lines against NextLat's per-domain bars.
Interactive visualization companion: full viz page