This episode explores Lattice, a 2025 paper from Google Research and Google DeepMind that asks whether a Transformer’s growing key-value cache can be compressed into a fixed set of memory slots without losing the long-context behavior users care about. It explains why this matters by contrasting standard attention’s unbounded cache with linear attention, recurrent state models, and fast-weight associative memory, framing the problem as memory compression rather than a rejection of Transformers. The discussion focuses on Lattice’s core idea: treat memory as an online low-rank factorization, reconstruct each new token from the current slots, and write only the residual through a single gradient-style update whose gate and direction arise from the math. Listeners would find it interesting because it gets into the real tradeoff between elegant compression and practical accuracy, including whether learned fixed-slot memory can beat simpler industry tactics like quantizing, sharding, or evicting cache entries.
Sources:
1. Lattice: Learning to Efficiently Compress the Memory — Mahdi Karami, Razvan Pascanu, Vahab Mirrokni, 2025
http://arxiv.org/abs/2504.056462. Compressive Transformers for Long-Range Sequence Modelling — Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Timothy P. Lillicrap, 2019
https://scholar.google.com/scholar?q=Compressive+Transformers+for+Long-Range+Sequence+Modelling3. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, Francois Fleuret, 2020
https://scholar.google.com/scholar?q=Transformers+are+RNNs%3A+Fast+Autoregressive+Transformers+with+Linear+Attention4. Palu: Compressing KV-Cache with Low-Rank Projection — Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Chong-Yan Chen, Yu-Fang Hu, Pei-Shuo Wang, Ning-Chi Huang, Luis Ceze, Mohamed S. Abdelfattah, Kai-Chiang Wu, 2024
https://scholar.google.com/scholar?q=Palu%3A+Compressing+KV-Cache+with+Low-Rank+Projection5. Lattice: Learning to Efficiently Compress the Memory — Mahdi Karami, Razvan Pascanu, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=Lattice%3A+Learning+to+Efficiently+Compress+the+Memory6. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jurgen Schmidhuber, 2021
https://scholar.google.com/scholar?q=Linear+Transformers+Are+Secretly+Fast+Weight+Programmers7. Gated Delta Networks: Improving Mamba2 with Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh, 2024
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule8. Kimi Linear: An Expressive, Efficient Attention Architecture — Kimi Team; Yu Zhang, Zongyu Lin, Xingcheng Yao, Jiaxi Hu, Fanqing Meng, et al., 2025
https://scholar.google.com/scholar?q=Kimi+Linear%3A+An+Expressive%2C+Efficient+Attention+Architecture9. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, Yoon Kim, 2024
https://scholar.google.com/scholar?q=Parallelizing+Linear+Transformers+with+the+Delta+Rule+over+Sequence+Length10. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, et al., 2024
https://scholar.google.com/scholar?q=Learning+to+%28Learn+at+Test+Time%29%3A+RNNs+with+Expressive+Hidden+States11. Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff — Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, Christopher Re, 2024
https://scholar.google.com/scholar?q=Simple+Linear+Attention+Language+Models+Balance+the+Recall-Throughput+Tradeoff12. Test-time Regression: a Unifying Framework for Designing Sequence Models with Associative Memory — Ke Alexander Wang, Jiaxin Shi, Emily B. Fox, 2025
https://scholar.google.com/scholar?q=Test-time+Regression%3A+a+Unifying+Framework+for+Designing+Sequence+Models+with+Associative+Memory13. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://arxiv.org/abs/2502.1600214. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yihua Cheng et al., 2025
https://arxiv.org/abs/2510.0966515. No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization — June Yong Yang et al., 2024
https://arxiv.org/abs/2402.1809616. Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters — Zhiyu Guo et al., 2024
https://arxiv.org/abs/2406.1233517. ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification — Yefei He et al., 2024
https://arxiv.org/abs/2405.1425618. State-space Models can Learn In-Context by Gradient Descent — Neeraj Mohan Sushma et al., 2024
https://arxiv.org/abs/2410.1168719. Test-Time Training Done Right — Tianyuan Zhang et al., 2025
https://arxiv.org/abs/2505.2388420. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2025
https://arxiv.org/abs/2501.0066321. AI Post Transformers: TRELLIS and Bounded-Memory Transformer KV Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-trellis-and-bounded-memory-transformer-k-81f237.mp322. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp323. AI Post Transformers: Gated Delta Networks for Long-Context Retrieval — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-gated-delta-networks-for-long-context-re-706d85.mp324. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp325. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp326. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp327. AI Post Transformers: Long Context Pre-Training with Lighthouse Attention — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-long-context-pre-training-with-lighthous-e85bbe.mp328. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3Interactive Visualization: Lattice: Fixed-Slot Compression for Transformer Memory