arXiv 2504.05646 Google Research + Google DeepMind Bounded-memory transformer state

Lattice: Fixed-Slot Compression for Transformer Memory

This companion page visualizes the paper’s central trade: replace an ever-growing Transformer KV cache with a fixed bank of memory slots, reconstruct each incoming token from that bank, and write back only the residual that still looks new. The page is about geometry, budgets, and system tradeoffs, not prose recap.

1 Steepest-descent-style write step per token.
M slots Memory budget set by slot count, not context length.
110M → 340M Model scales called out in the episode discussion.
2k → 16k Context slices explicitly discussed for Pile and Books3.

Memory Pressure

Standard attention stores an addressable trace of every past token. Lattice pays approximation instead of cache growth by forcing the history into a fixed slot bank.

Full KV Quantized KV Lattice Linear recurrent state

Bounded-memory lens

Hover the cache cells and frontier lines. The upper half contrasts token-addressable storage with fixed-slot compression; the lower chart shows why decode becomes memory-bound.
The page treats Lattice as a bridge design: still transformer-shaped in spirit, but borrowing recurrent and fast-weight ideas to keep state size fixed.

Orthogonal Write Rule

Step through four token situations. The update is small when the current slots already span the token, and larger when a genuinely new direction survives reconstruction.

Step focus

Bright cells in the delta map mark where the slot bank actually moves. Repetition and revision are both possible, but exact token overwrite is no longer free.

Benchmarks and Budget Curves

Exact Lattice values stated in the transcript anchor the loss view. The surrounding comparison numbers and matched-budget curves are stylized mock data for visual intuition.

Task texture

The side heatmap is a stylized six-task view for the larger model. It shows the episode’s narrower claim: promising language-modeling evidence, not yet a full systems sweep.

Deployment Tradeoffs

Different operators care about different frontiers. A serve-heavy stack may prefer cache engineering; an edge or streaming stack may prefer fixed-state memory even with softer exact recall.

Mode verdict

References

1. Lattice: Learning to Efficiently Compress the Memory (2025), Mahdi Karami, Razvan Pascanu, Vahab Mirrokni.
3. Transformers are RNNs (2020), Angelos Katharopoulos et al.
6. Gated Delta Networks (2024), Songlin Yang, Jan Kautz, Ali Hatamizadeh.
7. Titans: Learning to Memorize at Test Time (2025), Ali Behrouz, Peilin Zhong, Vahab Mirrokni.
8. No Token Left Behind (2024), June Yong Yang et al.
9. ZipCache (2024), Yefei He et al.
10. KVLink (2025), Jingbo Yang et al.