Visual companion page for the podcast episode. The core idea: let a standard autoregressive transformer mutate a small subset of its own weights during inference, using a next-token-prediction-aligned update signal.
The visual below contrasts ordinary transformer serving with in-place test-time training. Click modes to see where memory lives: tokens/KV cache versus temporary writable parameters.
Step through one transformer block. The writable component is the final MLP projection. The paper’s second move is to use an autoregressive, next-token-aligned adaptation objective instead of a generic reconstruction loss.
These illustrative charts map the discussion in the episode: promising long-context improvements, but not a free retrofit. Toggle the comparison view and hover points for detail.
This heatmap compares four adaptation routes across six deployment criteria. In-place TTT competes with longer context, retrieval, and external memory by shifting part of the working memory into weights.