arXiv:2604.06169 ByteDance Seed + Peking University Posted Apr 7, 2026

In-Place Test-Time Training
for Transformers

Visual companion page for the podcast episode. The core idea: let a standard autoregressive transformer mutate a small subset of its own weights during inference, using a next-token-prediction-aligned update signal.

Fast weights
MLP final projection
Serving mode shift
Forward → Update → Forward
Main skepticism
Architecture drop-in ≠ training-free

From static inference to dynamic inference

The visual below contrasts ordinary transformer serving with in-place test-time training. Click modes to see where memory lives: tokens/KV cache versus temporary writable parameters.

token / context flow
standard transformer compute
fast-weight updates
latency / systems overhead

What gets updated, and when?

Step through one transformer block. The writable component is the final MLP projection. The paper’s second move is to use an autoregressive, next-token-aligned adaptation objective instead of a generic reconstruction loss.

updated fast-weight matrix cells
high adaptation pressure
aligned prediction signal

Mocked result surfaces: gains, cost, and skepticism

These illustrative charts map the discussion in the episode: promising long-context improvements, but not a free retrofit. Toggle the comparison view and hover points for detail.

Where is the memory stored?

This heatmap compares four adaptation routes across six deployment criteria. In-place TTT competes with longer context, retrieval, and external memory by shifting part of the working memory into weights.