AI Post Transformers · Episode Companion

In-Place Test-Time Training Turns Fast Weights Into Online Memory

Guhao Feng, Shengjie Luo (equal contribution), Kai Hua, Ge Zhang, Di He, Wenhao Huang, Tianle Cai — ByteDance Seed, Peking University, Anthropic
arXiv:2604.06169 Submitted April 7, 2026 Qwen3-4B · LLaMA-3.1-8B · Qwen3-14B Interactive Viz ↗
A pretrained model's own gated-MLP down-projection becomes an online-updating fast weight — no new parameters, no architecture change — trained with an LM-aligned target instead of reconstruction, and applied via a causality-preserving context-parallel scan.

The Gated MLP, Repurposed

Standard gated MLP: out = (act(W_gate·x) · W_up·x) · W_down. In-Place TTT freezes W_gate and W_up as slow weights — only W_down moves, doing double duty as both a working forward-pass matrix and the thing being trained online, one chunk at a time.

Step Through: W_down as Evolving Fast Weight

Each chunk of tokens triggers an apply then an update. Click through chunks to watch the fast-weight matrix state drift from its pretrained initialization.

Chunk 0 — pretrained W_down, no updates applied yet.

Reconstruction Target vs. LM-Aligned Target

Prior TTT work (Sun et al. 2024) sets the training value to roughly the token's own embedding. In-Place TTT instead builds the value from a causal 1D convolution over embeddings, run through a trainable projection — provably raising the correct next-token logit (Theorem 1), where the reconstruction target's expected effect is ~zero.

Context-Parallel Scan

The update's associative structure lets chunk deltas be computed independently, then aggregated with a single prefix-sum, then applied in parallel — mathematically equivalent to a strict sequential update. Causal padding in the convolution keeps chunk i's target blind to anything past chunk i.

RULER: Qwen3-4B-Base, Baseline vs. In-Place TTT

Baseline edges ahead at 4k (96.6 vs 96.1) — then the story flips. By 64k the gap is 78.7 vs 74.3, holding out to 256k, double the trained context window.

4k/64k/128k/256k values as reported in the paper. 8k/16k/32k interpolated for illustration — paper states only that these deltas are sub-1pt.

RULER-64k Gain Across Model Families

In-Place TTT score minus baseline, at 64k context.

Stacks with YaRN (Qwen3-4B)

RoPE extension composes rather than conflicts.

Efficiency: Barely Moves

Figure 4's prefill throughput and peak memory stay close to baseline across context lengths and backbones — the RULER gains aren't hiding a compute tax.

Illustrative reproduction of Figure 4's near-flat trend; exact figures not stated in the transcript.

Chunk Size: Why 512–1024 Beats 2048

Bigger chunks mean W_down updates less often. Too coarse and the fast-weight state goes stale between updates; too fine and parallelism collapses. The sweet spot is 512–1024, with 1024 winning on efficiency.

Illustrative curve shaped to the described trade-off; exact ablation numbers not stated in the transcript.

State Size Scaling

More TTT-enabled layers → monotonically better RULER scores.

Objective Ablation Heatmap

Dropping either component hurts — convolution matters more long-context, projection more short-context.

From-Scratch: Sliding-Window Perplexity (The Pile, 32k context)

500M & 1.5B scale, vs. GLA, DeltaNet, and LaCT on the same SWA backbone. In-Place TTT posts the lowest perplexity at both scales.

Illustrative relative ranking; exact perplexity values not stated in the transcript.

Which Deltas Are Signal, Which Are Noise?

No seeds, no error bars. Sub-1pt moves at 4k/8k/16k/32k are within single-run noise. 64k and 128k — four-point-plus gaps, replicated across two model families — look like a real pattern.

Big Effect vs. Small Effect (Table 3)

Commonsense benchmarks move by fractions of a point. Long-context RULER scores move by double digits. When it's real, it doesn't need a seed argument to be convincing.

Repurpose vs. Add: In-Place TTT vs. Titans

Titans bolts on a new memory module — more capacity, more retraining pressure. In-Place TTT reuses W_down, which prior work already showed functions as key-value memory. Cheaper and drop-in, but now one matrix does two jobs — and nobody checked interference with LoRA or ROME-style edits that also target down_proj.

References