The Gated MLP, Repurposed
Standard gated MLP: out = (act(W_gate·x) · W_up·x) · W_down. In-Place TTT freezes W_gate and W_up as slow weights — only W_down moves, doing double duty as both a working forward-pass matrix and the thing being trained online, one chunk at a time.
Step Through: W_down as Evolving Fast Weight
Each chunk of tokens triggers an apply then an update. Click through chunks to watch the fast-weight matrix state drift from its pretrained initialization.
Reconstruction Target vs. LM-Aligned Target
Prior TTT work (Sun et al. 2024) sets the training value to roughly the token's own embedding. In-Place TTT instead builds the value from a causal 1D convolution over embeddings, run through a trainable projection — provably raising the correct next-token logit (Theorem 1), where the reconstruction target's expected effect is ~zero.
Context-Parallel Scan
The update's associative structure lets chunk deltas be computed independently, then aggregated with a single prefix-sum, then applied in parallel — mathematically equivalent to a strict sequential update. Causal padding in the convolution keeps chunk i's target blind to anything past chunk i.
RULER: Qwen3-4B-Base, Baseline vs. In-Place TTT
Baseline edges ahead at 4k (96.6 vs 96.1) — then the story flips. By 64k the gap is 78.7 vs 74.3, holding out to 256k, double the trained context window.
RULER-64k Gain Across Model Families
In-Place TTT score minus baseline, at 64k context.
Stacks with YaRN (Qwen3-4B)
RoPE extension composes rather than conflicts.
Efficiency: Barely Moves
Figure 4's prefill throughput and peak memory stay close to baseline across context lengths and backbones — the RULER gains aren't hiding a compute tax.
Chunk Size: Why 512–1024 Beats 2048
Bigger chunks mean W_down updates less often. Too coarse and the fast-weight state goes stale between updates; too fine and parallelism collapses. The sweet spot is 512–1024, with 1024 winning on efficiency.
State Size Scaling
More TTT-enabled layers → monotonically better RULER scores.
Objective Ablation Heatmap
Dropping either component hurts — convolution matters more long-context, projection more short-context.
From-Scratch: Sliding-Window Perplexity (The Pile, 32k context)
500M & 1.5B scale, vs. GLA, DeltaNet, and LaCT on the same SWA backbone. In-Place TTT posts the lowest perplexity at both scales.
Which Deltas Are Signal, Which Are Noise?
No seeds, no error bars. Sub-1pt moves at 4k/8k/16k/32k are within single-run noise. 64k and 128k — four-point-plus gaps, replicated across two model families — look like a real pattern.
Big Effect vs. Small Effect (Table 3)
Commonsense benchmarks move by fractions of a point. Long-context RULER scores move by double digits. When it's real, it doesn't need a seed argument to be convincing.
Repurpose vs. Add: In-Place TTT vs. Titans
Titans bolts on a new memory module — more capacity, more retraining pressure. In-Place TTT reuses W_down, which prior work already showed functions as key-value memory. Cheaper and drop-in, but now one matrix does two jobs — and nobody checked interference with LoRA or ROME-style edits that also target down_proj.