AI Post Transformers · Episode Companion

Tensor Cache: Compressing Evicted Tokens into Fixed-Size Memory

Swain, Han, Weidele, Martino, Torralba · 2026 MIT · U. Toronto · IBM Research arXiv:2605.22884 ↗

Deletion vs. Compression

Sliding-window attention caps memory by throwing evicted tokens away for good. Tensor Cache keeps the same fixed-size local window, but folds every evicted key/value pair into a second, fixed-size matrix instead of discarding it. Toggle the two regimes below to see where the evicted token actually goes.

Two-Tier Layout, Per Layer, Per Head

Every attention head carries both structures side by side. L1 does exact softmax attention over the last W tokens. L2 is a fixed D×D matrix that absorbs whatever L1 evicts, read out with a single matrix multiply rather than a token-by-token scan.

The Write Rule

Fast-weight memory (Schmidhuber, 1992) plus the Schlag et al. (2021) linear-attention identity: a query against the compressed matrix collapses to a dot product times a value — one matmul stands in for attention over everything superposed inside it.

A ← λ·A + η·(k ⊗ v)
step 0 / 4 — matrix empty

The Chunk-Mean Bug

Batching training by averaging keys and values per chunk before taking the outer product looks harmless — until the cross terms where key index ≠ value index start polluting memory. A chunk of length C bakes in C²−C spurious associations that never occur during real streaming inference.

Language-Modeling Quality vs. Context Length

OpenWebText (130M params) and Shakespeare replication, converged regime. Lower NLL is better. Click a legend entry to isolate a method. At 32,768 tokens: Tensor Cache 5.14, Full KV 6.00.

Retained State vs. Context Length

Every bounded method sits flat near 0.75–0.80 GB from 4K out to 128K tokens. Full KV grows linearly, reaching ~7.75 GB at 128K.

The Training Tax

Inference-time efficiency is bought with a one-time training cost. At gap 1024, Tensor Cache is ~324× slower to train than Full KV and uses ~10× more training memory.

Capacity Ceiling of Matrix A

Writes tolerated before the oldest association's reconstruction error blows past threshold. Decay trades capacity for staying fresh: at D=128, no-decay holds 240 writes, the actual trained decay setting holds ~84.

Write-Rule Ablation

An antisymmetric wedge-product write (direction should matter, in theory) underperforms both the plain outer product and the delta rule. Reported as a negative result.

Matched-Gap Recall: Accuracy Held, State Cut

Tensor Cache matches Full KV's near-perfect recall at every gap while retaining only 16–28% of the state Full KV would keep. Hover a cell for the exact value.

What The Paper Doesn't Show Yet

Flagged in the episode discussion — none of these break the mechanism, but none are resolved by the current experiments either.

Max scale: 130M params, no pretrained checkpoint 324× training-time tax vs. Full KV at gap 1024 Single-seed beyond 4K context (σ=0.66 @ L=1024) H2O / SnapKV / CAOTE / LESS absent from main tables Loses exact-entry addressability (breaks CacheGen-style reuse) No interpretability study on the learned gate

References