← All episodes Training Transformers for Compressible KV Caches

Training Transformers for Compressible KV Caches

Sep 30, 2026
This episode explores a paper arguing that KV cache compressibility is not an inherent property of an input, but a structural property of how a transformer's weights happen to represent a computation. Using a histogram-counting thought experiment, the authors prove that two transformers can compute the identical function while one is trivially compressible and the other is fundamentally not, exposing a blind spot in existing inference-time compression techniques like heavy-hitter eviction, attention-sink methods, and optimization-based approaches such as Cartridges and Attention Matching. The discussion surveys the broader landscape of memory-saving strategies, contrasting architectural fixes like linear attention and state space models against post-hoc interventions on standard transformers, and explains why long-context and agentic workloads make cache size a serious bottleneck. The hosts debate whether the paper's formal construction is just a toy example or a meaningful warning, concluding that it reframes compressibility as something that must be deliberately trained for rather than assumed to emerge naturally from ordinary pretraining. Listeners interested in efficient LLM serving will find this useful for understanding why some models resist compression no matter how sophisticated the algorithm applied to them.
Sources:
1. Training Transformers for KV Cache Compressibility — Yoav Gelberg, Yam Eitan, Michael Bronstein, Yarin Gal, Haggai Maron, 2026
http://arxiv.org/abs/2605.05971
2. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, Sumit Sanghai, 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, Beidi Chen, 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. Efficient Streaming Language Models with Attention Sinks (StreamingLLM) — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks+%28StreamingLLM%29
5. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI (Aixin Liu et al.), 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
6. Cartridges: Lightweight and general-purpose long context representations via self-study — Sabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha, Dylan Zinsley, Emily Liu, Will Tennien, Atri Rudra, James Zou, Azalia Mirhoseini, Christopher Ré, 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+general-purpose+long+context+representations+via+self-study
7. Fast KV compaction via attention matching — Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim, 2026
https://scholar.google.com/scholar?q=Fast+KV+compaction+via+attention+matching
8. Dynamic chunking for end-to-end hierarchical sequence modeling — Sukjun Hwang, Brandon Wang, Albert Gu, 2025
https://scholar.google.com/scholar?q=Dynamic+chunking+for+end-to-end+hierarchical+sequence+modeling
9. DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads — Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, Song Han, 2025
https://scholar.google.com/scholar?q=DuoAttention%3A+Efficient+long-context+LLM+inference+with+retrieval+and+streaming+heads
10. Lexico: Extreme KV cache compression via sparse coding over universal dictionaries — Junhyuck Kim, Jongho Park, Jaewoong Cho, Dimitris Papailiopoulos, 2025
https://scholar.google.com/scholar?q=Lexico%3A+Extreme+KV+cache+compression+via+sparse+coding+over+universal+dictionaries
Interactive Visualization: Training Transformers for Compressible KV Caches