1 00:00:01,000 --> 00:00:55,009 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called In-Place Test-Time Training, from Guhao Feng and Shengjie Luo — equal contribution — plus five more co-authors: Kai Hua, Ge Zhang, Di He, Wenhao Huang, and Tianle Cai. So Feng and Luo, et al., seven authors total, out of ByteDance Seed, Peking University, and Anthropic. It went up on arXiv on April 7th, 2026. Here's the claim worth sitting with: this paper wants to give an already-trained model the ability to keep learning while it's running, without touching its architecture. Not fine-tuning. Not a new layer bolted on top. It wants to repurpose parts the model already has. 2 00:00:55,009 --> 00:01:04,158 [Hal Turing] So let's get into the actual mechanics, Ada. When you say the MLP block becomes the fast weights, what's literally changing during a forward pass? 3 00:01:04,158 --> 00:01:54,917 [Dr. Ada Shannon] Picture the standard gated MLP: output equals the gate-activated projection times the up-projection, times the down-projection. Three matrices — W-gate, W-up, W-down. In-Place TTT freezes W-gate and W-up entirely; those stay slow weights. W-down is the one that moves. For each chunk of tokens there's an apply-then-update cycle: the current W-down state first processes that chunk's activations to produce the output, exactly like a normal MLP. Then, using those same activations as keys and a derived target as values, W-down takes one gradient step and updates itself before the next chunk arrives. So it's doing double duty — a working part of the forward pass every chunk, and also the thing being trained online. 4 00:01:54,917 --> 00:02:03,462 [Hal Turing] And that target — that's where the LM-aligned piece comes in? Because just associating a token with itself doesn't sound like it teaches much. 5 00:02:03,462 --> 00:02:44,329 [Dr. Ada Shannon] That's the gap they close. Prior TTT work sets the value to roughly the current token's own embedding — a reconstruction target. This paper instead builds the target as a causal 1D convolution over token embeddings, run through a trainable projection matrix, so it can pull in information from upcoming tokens. The fast weight gets trained to store next-token-predictive content, not just memorize what it already saw. They back this with Theorem 1, an induction-head analysis: the LM-aligned target provably raises the logit of the correct next token, while the reconstruction target's expected effect on that same logit is essentially zero. That's a proven result, not a hunch. 6 00:02:44,329 --> 00:02:52,920 [Hal Turing] Oh wait wait wait — hold on, doesn't pulling in future-token information wreck causality once you try to parallelize this across a whole context window? 7 00:02:52,920 --> 00:03:27,053 [Dr. Ada Shannon] That's the elegant part. The update's associative structure lets them run a context-parallel scan: compute each chunk's weight delta independently, then a single prefix-sum aggregates the deltas up to every chunk boundary, then everything applies in parallel. Causality holds because the convolution builds each target with causal padding, so chunk i's target never actually sees past chunk i. At document boundaries they reset W-down to its pretrained value so nothing leaks across sequences. It's mathematically equivalent to a strict sequential update, just computed all at once. 8 00:03:27,053 --> 00:03:32,440 [Hal Turing] Let's get to the drop-in numbers, then. What happened when they bolted this onto Qwen3-4B-Base? 9 00:03:32,440 --> 00:04:17,812 [Dr. Ada Shannon] On RULER, 4k through 256k, the baseline actually edges out In-Place TTT at 4k — 96.6 versus 96.1. But from there it flips: by 64k it's 78.7 versus 74.3, at 128k 77.0 versus 74.8, and extrapolating to 256k — double the trained window — it still holds, 43.9 versus 41.7. So a small dip at short context, then a widening advantage the longer the sequence gets. They repeated the recipe on LLaMA-3.1-8B and Qwen3-14B-Base: plus 2.1 at 64k for LLaMA, plus 2.7 at 64k for Qwen3-14B, and the gains survive stacking with YaRN for RoPE extension, 81.3 up to 82.5. 10 00:04:17,812 --> 00:04:23,060 [Hal Turing] And the from-scratch comparisons against the other TTT-style methods — how'd it stack up there? 11 00:04:23,060 --> 00:05:15,212 [Dr. Ada Shannon] At 500M and 1.5B scale on the Pile, 32k context, they compare against sliding-window attention, Gated Linear Attention, DeltaNet, and LaCT — Behrouz's Large Chunk TTT — with In-Place TTT and LaCT both built on the same SWA backbone for fairness. In-Place TTT posts the lowest sliding-window perplexity at both scales and keeps improving out to 32k. At 4B, trained from scratch for 120 billion tokens, the common-sense benchmarks — HellaSwag, ARC, MMLU, PIQA — barely move. The long-context numbers move a lot more: RULER-16k goes from 6.58 to 19.99 on the full-attention backbone, RULER-8k goes from 9.91 to 26.80 on SWA. 12 00:05:15,212 --> 00:05:26,822 [Hal Turing] Wait, I want to push on the chunk-size ablation for a second — bigger chunks should always win on a GPU, more parallelism. Why would 512 to 1024 beat 2048? 13 00:05:26,822 --> 00:05:55,568 [Dr. Ada Shannon] No, I actually disagree with that instinct, Hal. Bigger chunks mean W-down updates less often — you're not just trading efficiency for performance, you're changing how much the fast weights get to adapt within a stretch of context. Too coarse and the state goes stale between updates; too fine and you lose the parallelism that makes this practical at all. Their ablation puts the sweet spot at 512 to 1024, with 1024 winning on efficiency. It's a real trade-off curve. 14 00:05:55,568 --> 00:06:01,141 [Hal Turing] Okay, fair — I was treating it as a pure hardware question and ignoring that update frequency is doing real work. 15 00:06:01,141 --> 00:06:36,668 [Dr. Ada Shannon] Right. State size behaves the way you'd expect too — more TTT-enabled layers, more capacity, monotonically better RULER scores. They also strip the objective apart: drop either the convolution or the projection matrix and performance drops, confirming both earn their keep, convolution mattering more at long context, the projection more at short. And on the efficiency side, Figure 4 shows prefill throughput and peak memory for both SWA and full-attention backbones barely move at 8k, 32k, or 128k — whatever's driving those RULER gains isn't coming from a hidden compute tax. 16 00:06:36,668 --> 00:07:06,157 [Hal Turing] Let's get skeptical for a bit, Ada. In Table 1, Qwen3-4B-Base with In-Place TTT actually loses to the baseline at 4k context — 96.1 versus 96.6 — and most of the other length gains are under a point. Only 64k and 128k really separate from baseline, and there's no mention of seeds or variance anywhere. How much of this 'widening advantage' story survives someone just asking whether it's noise? 17 00:07:06,157 --> 00:07:43,402 [Dr. Ada Shannon] Some of it doesn't survive that question as written. A half-point dip at 4k and sub-one-point moves at 8k, 16k, and 32k are well within what you'd expect from one training run with no seed averaging. Where I won't call it noise is 64k and 128k — 78.7 versus 74.3, and 77.0 versus 74.8, four-point-plus gaps in the same direction across two different lengths and two separate model families in Table 2. That's a pattern. They just should have run it three times and shown error bars instead of making listeners do this math themselves. 18 00:07:43,402 --> 00:08:13,588 [Hal Turing] Fair — I was treating length as one continuous story instead of splitting signal from noise. But here's what actually bugs me more: the only place they race In-Place TTT against LaCT, GLA, and DeltaNet head to head is the 500M and 1.5B from-scratch runs in Figure 2. The 4B-to-14B drop-in results — the actual headline claim — only compare against the unmodified baseline. LaCT never shows up at that scale. 19 00:08:13,588 --> 00:08:44,052 [Dr. Ada Shannon] And that's not a minor omission, Hal, that's the paper eating its own argument. LaCT is their closest competitor, same chunk-wise philosophy, and it gets tested exactly where it's weakest for them and skipped exactly where their headline claim lives. You can't validate 'we beat other TTT methods' at 1.5B and 'we work as a drop-in' at 14B separately, then splice them into one abstract sentence like it's a single experiment. Nobody ever ran LaCT as a drop-in on Qwen3-14B. 20 00:08:44,052 --> 00:09:00,167 [Hal Turing] I'll grant the gap, but isn't some of that just compute budget? Continually training three competing methods on a 14B model, plus YaRN extension for every baseline, is a lot of GPU-hours to spend on one ablation table nobody's forcing them to run. 21 00:09:00,167 --> 00:09:33,604 [Dr. Ada Shannon] Could be. I'm not accusing them of burying a losing number. But budget constraints don't change what the evidence actually supports — if you can't afford the comparison, you don't get to imply it would go your way. The honest version of this paper says 'we beat LaCT small, and drop-in gains hold large, but we never tested whether LaCT drops in well large.' That's a real result. It's just smaller than 'consistently outperforms competitive TTT-related approaches,' which is the actual line in their abstract. 22 00:09:33,604 --> 00:09:54,827 [Hal Turing] There's a similar gap between narrative and evidence in the intro itself. They motivate all of this with 'continuous streams of new information' and learning from 'unbounded streams of experience like humans' — that's Silver and Sutton's framing, the 'Era of Experience' piece out of Google DeepMind, 2025. But then in the implementation section— 23 00:09:54,827 --> 00:10:39,920 [Dr. Ada Shannon] —wait, hold on, I caught that too. Section 3.4: fast weights reset to their pretrained state at every document boundary. Zero persistence across documents, zero across sessions. So whatever this is, it isn't continual learning in the sense the intro promises — it's excellent intra-context retrieval, up to 256k tokens, inside a single forward pass. Same story with Table 3's commonsense numbers: HellaSwag up 0.18, ARC-E up 0.46, no seeds mentioned — that's noise. Compare it to RULER-16k jumping from 6.58 to 19.99 in that same table. When an effect is real, it doesn't need a seed argument to be convincing. 24 00:10:39,920 --> 00:10:57,660 [Hal Turing] So structurally, how does this compare to Titans — the Behrouz, Zhong, and Mirrokni paper out of Google Research, 2024? That's the other big test-time-memory architecture people cite in the same breath as this one, and it takes a pretty different bet on where the memory should live. 25 00:10:57,660 --> 00:11:58,821 [Dr. Ada Shannon] Titans bolts on a brand-new memory module alongside attention — an 'add' philosophy: more capacity, more retraining pressure. In-Place TTT is 'repurpose' — no new parameters, reuse W_down, which Geva, Schuster, Berant, and Levy showed functions as key-value memory, out of Bar-Ilan University, back in 2020. Repurposing is cheaper and drop-in by construction, but now one matrix does two jobs. It never asks the obvious follow-up either: down_proj is a favorite LoRA target, per Hu and colleagues' paper out of Microsoft, 2022, and it's where ROME-style edits — Meng, Bau, Andonian, and Belinkov, out of MIT and Northeastern, 2022 — get written. Overwriting it per request without checking for interference is a real gap. This whole line traces back to Sun, Li, Dalal, and colleagues' 2024 paper out of Stanford — the two barriers named here are exactly what that paper left open. 26 00:11:58,821 --> 00:12:18,697 [Hal Turing] So practically, what does a team actually do with this? Sounds like: real signal at long context on the specific checkpoints tested, but verify on your own model before trusting it, don't assume it composes cleanly with LoRA adapters or ROME-style edited weights, and don't sell it internally as cross-session memory, because it isn't that. 27 00:12:18,697 --> 00:12:51,530 [Dr. Ada Shannon] The paper itself flags trying other loss functions and optimizers, and integrating with other efficient backbones like GLA or SSMs, since those have MLP blocks too. What the evidence gaps point to is more specific: seeded runs so we know which numbers are real, and a LaCT or DeltaNet comparison run as an actual drop-in at 4B-plus. Until that experiment exists, 'consistently outperforms competitive TTT-related approaches' is a claim about small models, not the one being sold. 28 00:12:51,530 --> 00:13:20,602 [Hal Turing] Repurposing W_down instead of bolting on new memory is genuinely clever and cheap, and the long-context gains at 64k-plus look real. But 'continual learning from unbounded experience' oversells what's actually a strong retrieval trick, and the 'beats other methods' claim was never tested at the scale it's sold at. Worth watching, not worth trusting blindly yet. That's it for this one — thanks for listening, and we'll catch you next time.