1 00:00:01,000 --> 00:00:37,250 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're covering Next-Latent Prediction Transformers Learn Compact World Models, by Jayden Teoh and nine co-authors, all at Microsoft Research. The arXiv listing is on its fourth revision, dated June 2026. The claim is that a small auxiliary loss can make an ordinary transformer squeeze its history into a compact state, without touching the architecture. Ada, the paper opens with Ptolemy versus Copernicus. Why start with astronomy? 2 00:00:37,250 --> 00:01:16,849 [Dr. Ada Shannon] Because both models predicted the sky well, and only one generalized. Ptolemy's had the Moon swinging to twice as close as at other times. That's a model built to fit observations, not to explain them. The authors say transformers have the same problem. Recurrent networks carry a fixed-size state and apply the same update rule at every step, so compression comes for free. A transformer replaces that with a KV memory that grows with the sequence, and attention can look up any old token at will. Nothing forces it to keep a compact summary. So it learns shortcuts that fit the training data and break outside it. Same team, by the way, that built the Belief State Transformer and Joint Multi-Token Prediction. This is a follow-up to their own line of work. 3 00:01:16,849 --> 00:01:23,599 [Hal Turing] Okay, I keep hearing "belief state." What is that, concretely, and why should a transformer care? 4 00:01:23,599 --> 00:02:01,824 [Dr. Ada Shannon] It's a sufficient statistic of the past. Kaelbling, Littman and Cassandra at Brown formalized it for POMDPs in 1998, and Striebel had the same idea as an information state in 1965. Think of a taxi. Its current intersection tells you everything about where it can go next. The whole route so far adds nothing. Formally, a belief state predicts the future exactly as well as the full history does. Now the failure case. Vafa and colleagues at Harvard, 2024, trained transformers on Manhattan taxi routes. The models hit 100 percent next-turn accuracy, but the street maps you could reconstruct from their internals were incoherent, with streets that don't— 5 00:02:01,824 --> 00:02:08,774 [Hal Turing] Wait wait wait, a hundred percent accurate with a nonsense map? How does that model drive anywhere? 6 00:02:08,774 --> 00:02:52,974 [Dr. Ada Shannon] It drives fine until you close a street. Then the shortcut collapses. That's the point: next-token accuracy can't tell you whether there's a real world model inside. And the training objective is myopic. Bachmann and Nagarajan, out of EPFL and Carnegie Mellon in 2024, showed the Clever Hans cheat. The model gets earlier tokens right by exploiting local patterns, so it never learns the hard planning step. Their Path-Star graph task exposes it. Then the Belief State Transformer paper by Hu and colleagues at Microsoft Research, 2025, proved in its Theorem 3 that next-token consistency alone doesn't guarantee a belief state. One vocabulary note. A latent state is anything in the residual stream. The hidden state is specifically the final-layer, pre-logit vector at a position. NextLat only touches that one. 7 00:02:52,974 --> 00:02:57,874 [Hal Turing] So the same group already tried fixing this twice. What went wrong? 8 00:02:57,874 --> 00:03:36,774 [Dr. Ada Shannon] Cost, and a condition. The Belief State Transformer runs a forward encoder and a backward encoder over every prefix and suffix pair. That gives quadratic training signal and the belief-state guarantee, at more than double the parameters. JTP, by Ahn, Lamb and Langford in 2025, predicts several tokens jointly and is much cheaper. But its guarantee depends on k-observability: a system is k-observable if the next k tokens' distribution pins down the entire future. JTP only gets belief states when its horizon d is at least k minus one. So you get an expensive guarantee, or a cheap one that depends on a number you don't know. 9 00:03:36,774 --> 00:03:39,674 [Hal Turing] And NextLat gets out of that how? 10 00:03:39,674 --> 00:04:15,549 [Dr. Ada Shannon] By borrowing from RL. MuZero at DeepMind in 2020, Dreamer from Hafner, and Genie all learn a dynamics model: state plus action gives next state. Tang and colleagues at DeepMind, 2023, showed you can train a representation to predict its own future latent, with a stop-gradient on the target so it doesn't collapse. NextLat treats the next token as the action and the hidden state as the latent. A small network predicts the transformer's next hidden state from the current one plus the token. It's used only in training. The only other latent-space objective in language modeling is LLM-JEPA, and it needs paired data. 11 00:04:15,549 --> 00:04:24,250 [Hal Turing] You mentioned a speed payoff. I know speculative decoding is draft-and-verify, but walk me through the self-speculative part. 12 00:04:24,250 --> 00:04:49,024 [Dr. Ada Shannon] Leviathan and colleagues at Google, 2023: a cheap draft proposes tokens, and the big model checks them all in one parallel pass. Self-speculative means the same model drafts, with no second network. Multi-token prediction, from Gloeckle and colleagues at Meta in 2024, does that with extra heads, so the draft length is frozen at the training horizon. NextLat's dynamics model can be rolled forward recursively, so draft length isn't fixed. We'll get to the numbers later. 13 00:04:49,024 --> 00:05:15,774 [Hal Turing] So the paper makes three claims. Theory that the latents converge to belief states. A loss that leaves architecture, parallel training and inference untouched. And gains across world modeling, planning, reasoning and language modeling, plus that flexible-length drafting. One thing I'll say plainly: measuring compression on a model that already scores perfectly on next-turn accuracy is the right test. Accuracy was never going to separate these models. 14 00:05:15,774 --> 00:05:24,424 [Hal Turing] Okay, Ada, the machinery. There's a theorem at the center of this. What does it promise, and what has to hold for it to be true? 15 00:05:24,424 --> 00:06:19,849 [Dr. Ada Shannon] It promises that h t is a belief state, given two idealized conditions on three parts: the transformer, the output head, and the latent dynamics model. Next-token consistency says the head reading h t gives the true next-token distribution. Transition consistency says the dynamics model, given h t and the next token, reproduces the true law of the transformer's own next hidden state. If both hold, you can decode a token, update the state, and repeat to the end without ever seeing the prefix. The proof is backward induction. The last step follows from the first condition, and if the next state is a belief state, decode-then-update makes the current one a belief state too. It holds already at horizon one. Longer horizons only add signal, because the next hidden state encodes a whole distribution over the following token, not a one-hot label. The trained loss is cross-entropy plus a Smooth-L1 regression of the dynamics model's rolled-out states against the real hidden states, with a stop-gradient on the target, plus a— 16 00:06:19,849 --> 00:06:31,249 [Hal Turing] Sorry, hold on, that's what I don't get. The target is the model's own state. Why doesn't everything collapse to a constant, where predicting yourself is trivially perfect? 17 00:06:31,249 --> 00:07:17,324 [Dr. Ada Shannon] Two guards. The stop-gradient stops the target chasing the prediction, and cross-entropy forces h t to keep token information. The ablations say the stop-gradient matters most on Countdown and Path-Star. The third term is a KL: the frozen output head's distribution at the predicted state is pulled toward its distribution at the true state. It's distillation-like, and it plays the role of observation reconstruction in self-predictive RL. The dynamics model is a three-layer GELU MLP on the layer-normalized concatenation of h t and the next-token embedding, outputting a residual delta. It's discarded at inference unless you draft with it. At horizon one, training runs at 3.09 steps per second, the same as GPT. BST manages 0.89, with 2.57 billion parameters. NextLat is 1.40 billion in training and 1.32 at inference. 18 00:07:17,324 --> 00:07:24,974 [Hal Turing] Straight cost accounting, I respect that. Now show me results. Manhattan first, since BST couldn't afford it. 19 00:07:24,974 --> 00:08:25,499 [Dr. Ada Shannon] All four models score 100 percent on next-turn legality, so that test is saturated. Valid out-of-distribution trajectories: 98.7 percent against GPT's 97.0. Effective rank is the exponentiated entropy of the hidden states' normalized singular values, and lower means more compact. NextLat gets 52.7, GPT 160.1, JTP 215.8. Sequence compression asks whether two routes to one intersection get the same continuation: 0.71 versus 0.65. Detour robustness ties MTP at 95 percent, with GPT at 85. The reconstructed map shows sparse, mostly local errors. On Countdown, NextLat scores 54.8, 57.6 and 58.7 at horizons one, four and eight. MTP gets 39.2, 49.7, 57.3. JTP gets 39.0, 49.4, 55.0. BST gets 42.3 and GPT 33.1, over three seeds. 20 00:08:25,499 --> 00:08:36,399 [Hal Turing] Wait, horizon one beating MTP and JTP by fifteen points? At one step they're barely doing anything extra. What is NextLat getting that they aren't? 21 00:08:36,399 --> 00:09:17,524 [Dr. Ada Shannon] The authors say lookahead. Failures concentrate in the final equation, what Ye and colleagues in 2025 called the regretful compromise, where the model corners itself and forces a wrong last step. On Path-Star, from Bachmann and Nagarajan in 2024, NextLat is near 100 percent on G(2,10), G(5,5) and G(7,7). BST solves the first two and fails G(7,7). That's the harder original setup, with 200,000 fixed samples, so it isn't comparable to the BST and JTP papers' numbers. On TinyStories, linear probes on frozen states show NextLat matching GPT at next token and leading at longer offsets. BST, MTP and JTP hurt next-token quality. 22 00:09:17,524 --> 00:09:28,899 [Hal Turing] Running the harder original setup and saying so is the kind of thing I like to see. So, the 1.3-billion-parameter run. Did it move actual language modeling? 23 00:09:28,899 --> 00:09:57,099 [Dr. Ada Shannon] Modestly. On 100 billion FineWeb-Edu tokens, NextLat at horizon two averages 59.21 zero-shot accuracy against GPT's 58.82, and the authors call that inconsistent across tasks. The firmer result is perplexity: 10.83 and 10.88 against GPT's 10.52, while JTP is 11.08 or worse and MTP 10.90 or worse. Horizon four for MTP and JTP doesn't change the accuracy picture. 24 00:09:57,099 --> 00:10:00,074 [Hal Turing] So where's the payoff? Drafting? 25 00:10:00,074 --> 00:10:39,374 [Dr. Ada Shannon] Yes. The MLP recurses: from the state and a drafted token it produces the next state, the head decodes the next token, and it repeats for as long as you like. Even trained at horizon one, it averages 3.5 accepted tokens on Wikipedia. At horizon two the speedups are 3.21x on Wikipedia, 3.32x on Books, 2.38x on Code and 2.87x on Math. JTP is about 1.9x and MTP about 1.7x. Draft length was picked per domain between two and ten, static, on eight B200s. JTP and MTP at horizon four close some of the gap and beat it on Code. 26 00:10:39,374 --> 00:10:51,024 [Hal Turing] Last one, the A5 experiment. A transformer that can't do the task trains an RNN that can. The student outgrew the teacher on the one exam the teacher failed. 27 00:10:51,024 --> 00:11:47,149 [Dr. Ada Shannon] That's the result. A5 is a word problem over even permutations of five elements, NC1-complete, out of reach for constant-depth transformers per Merrill and Sabharwal at NYU in 2023. They train a two-layer transformer on 12 tokens at horizon one, regression only. The transformer fails at 36 tokens. The co-trained MLP, 2.62 million parameters against 6.43 million, starts from the transformer's first-token state and clears 95 percent at 36 tokens. GPT trained directly on 36 still fails. Whether a shallow transformer can co-train an NC1 recurrence is posed as an open question. There's no truncated gradient through the transformer, so no horizon-dependent bias, and the sequential cost is d steps, not T. And unlike SSMs or hybrids, it's an objective change, not an architecture change. 28 00:11:47,149 --> 00:12:09,299 [Hal Turing] Ada, before we leave A5, I want a control. The MLP gets regressed onto a two-layer transformer's states, and that group has only sixty elements. Here's what I can't tell from the paper. Could the MLP just be learning group multiplication on its own, with the transformer acting as a very expensive label printer? What would settle it? 29 00:12:09,299 --> 00:12:45,049 [Dr. Ada Shannon] Nobody has ruled it out. The control is an MLP trained directly on state-and-token pairs, plus a small RNN trained on the task. The transformer limit doesn't cover this, and Merrill, Petty and Sabharwal showed in The Illusion of State in State-Space Models, 2025, that linear SSMs fail too. But a nonlinear recurrence isn't bound by either result. An MLP loop solving a sixty-state group is expected. It's one task, one length shift, regression-only, and the rollout starts from the transformer's first-token state. That's a good existence result. It is not evidence of NC1 computation. 30 00:12:45,049 --> 00:12:58,549 [Hal Turing] Parked as promising but unproven. Back to the theorem. Training minimizes a regression loss on a three-layer MLP, not exact equality. How close do the trained models get to the conditions? 31 00:12:58,549 --> 00:13:39,724 [Dr. Ada Shannon] Nobody measured it, and one appendix result points the wrong way. In Appendix E the Smooth-L1 latent loss rises during learning-rate cooldown, under both AdamW and Muon. That means transition consistency is getting worse at that stage. MSE, retuned coefficients and dropped stop-gradients didn't fix it. The authors point to five-token acceptance still improving. But acceptance is a token-space proxy scored against the model's own distribution, so it says little about the geometry of h. A scale-normalized latent error, or the same run without cooldown, would tell us. And nobody tests sufficiency directly: give h_t alone, without the attention context, and check whether next-k predictions survive. 32 00:13:39,724 --> 00:13:47,849 [Hal Turing] Hold on, that's actually— effective rank fifty-two versus a hundred sixty. Isn't that the sufficiency evidence? 33 00:13:47,849 --> 00:14:18,899 [Dr. Ada Shannon] It's compression. A rank-one state is very compact and useless. Theorem 3.2 doesn't force minimality either, because a state that copies raw history forward satisfies both equations. Shai and colleagues at Simplex found fractal belief-state geometry in the residual stream of ordinary next-token transformers on HMM data, in 2024. So probe a plain GPT with the same tools before crediting the loss. Rank is also measured only on Manhattan, from one run, and JTP's state is taken from a different place in the network. 34 00:14:18,899 --> 00:14:35,374 [Hal Turing] The Countdown gap at horizon one, 54.8 against about 39, is the cleanest win in the paper. But the claim is BST's guarantee at a fraction of the cost, and BST sits out Manhattan and the 1.3B run. 35 00:14:35,374 --> 00:15:02,149 [Dr. Ada Shannon] So the strongest baseline appears only where it loses. There's a second tension. Theory says the guarantee is independent of d, yet Countdown climbs to 58.7 at d equals 8 and acceptance improves with d. The guarantee is one thing. In practice, d still helps a lot. Also, the MTP paper itself reported that its gains show up at larger scales. So beating MTP at 1.3B means beating it where it's known to be weak. DeepSeek-V3's sequential MTP isn't compared at all. 36 00:15:02,149 --> 00:15:06,599 [Hal Turing] Then the 3.3x speedup, since that's the headline. 37 00:15:06,599 --> 00:15:46,974 [Dr. Ada Shannon] The speed win is real, but the accuracy edge is one seed across nine benchmarks, which is noise, and perplexity still loses to GPT. The draft length is the best static value per domain, swept from 2 to 10, and 'variable-length' is a capability, not a policy. There's no EAGLE, from Yuhui Li at Peking University in 2024, and no Medusa, from Tianle Cai at Princeton the same year. JTP at d equals 4 beats it on code, and batch size and temperature go unreported. The real strength is that the draft is a trained recurrence. Practically, try it as an auxiliary loss, and measure sufficiency yourself. The paper also never looks at the KV cache, or at using the dynamics model for planning. 38 00:15:46,974 --> 00:16:06,974 [Hal Turing] So my takeaway: a clean idea and a real speed win, resting on an idealized theorem and mostly proxy evidence. Before anyone calls these states belief states, the field needs direct sufficiency probes, error bounds, and BST matched at scale. Thanks for listening, everyone. See you next time.