← All episodes

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

Mar 25, 2026
This episode explores TurboQuant, a method for online vector quantization that aims to compress high-dimensional embeddings and transformer KV caches without any pretrained codebook or calibration pass. It explains how the paper connects classical rate-distortion theory to practical ML systems, contrasting mean-squared reconstruction error with inner-product preservation for tasks like retrieval and attention. The discussion highlights TurboQuant’s core idea of using random rotations and quantized Johnson-Lindenstrauss style sketches to make data-oblivious compression theoretically strong while still relevant to modern workloads. A listener would find it interesting because it probes whether a single, plug-and-play quantization scheme can approach information-theoretic limits while addressing real memory and bandwidth bottlenecks in large-scale AI systems.
Sources:
1. TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Amir Zandieh, Majid Daliri, Majid Hadian, Vahab Mirrokni, 2025
http://arxiv.org/abs/2504.19874
2. Vector Quantization — Robert M. Gray, 1984
https://scholar.google.com/scholar?q=Vector+Quantization
3. Product Quantization for Nearest Neighbor Search — Herve Jegou, Matthijs Douze, Cordelia Schmid, 2011
https://scholar.google.com/scholar?q=Product+Quantization+for+Nearest+Neighbor+Search
4. Optimized Product Quantization — Tiezheng Ge, Kaiming He, Qifa Ke, Jian Sun, 2013
https://scholar.google.com/scholar?q=Optimized+Product+Quantization
5. Additive Quantization for Extreme Vector Compression — Artem Babenko, Victor Lempitsky, 2014
https://scholar.google.com/scholar?q=Additive+Quantization+for+Extreme+Vector+Compression
6. Online Product Quantization — Donna Xu, Ivor W. Tsang, Ying Zhang, 2018
https://scholar.google.com/scholar?q=Online+Product+Quantization
7. Online Optimized Product Quantization — Chao Liu, Defu Lian, Min Nie, Huabin Xia, 2020
https://scholar.google.com/scholar?q=Online+Optimized+Product+Quantization
8. Locally-Adaptive Quantization for Streaming Vector Search — Cecilia Aguerrebere, Mark Hildebrand, Ishwar Singh Bhati, Theodore L. Willke, Mariano Tepper, 2024
https://scholar.google.com/scholar?q=Locally-Adaptive+Quantization+for+Streaming+Vector+Search
9. Real-time Indexing for Large-scale Recommendation by Streaming Vector Quantization Retriever — Xingyan Bin, Jianfei Cui, Wujie Yan, Zhichen Zhao, Xintian Han, Chongyang Yan, Feng Zhang, Xun Zhou, Qi Wu, Zuotao Liu, 2025
https://scholar.google.com/scholar?q=Real-time+Indexing+for+Large-scale+Recommendation+by+Streaming+Vector+Quantization+Retriever
10. Coding Theorems for a Discrete Source with a Fidelity Criterion — Claude E. Shannon, 1959
https://scholar.google.com/scholar?q=Coding+Theorems+for+a+Discrete+Source+with+a+Fidelity+Criterion
11. Rate Distortion Theory: A Mathematical Basis for Data Compression — Thomas Berger, 1971
https://scholar.google.com/scholar?q=Rate+Distortion+Theory%3A+A+Mathematical+Basis+for+Data+Compression
12. The Information Bottleneck Method — Naftali Tishby, Fernando C. Pereira, William Bialek, 1999
https://scholar.google.com/scholar?q=The+Information+Bottleneck+Method
13. End-to-end Optimized Image Compression — Johannes Balle, Valero Laparra, Eero P. Simoncelli, 2017
https://scholar.google.com/scholar?q=End-to-end+Optimized+Image+Compression
14. Similarity Estimation Techniques from Rounding Algorithms — Moses S. Charikar, 2002
https://scholar.google.com/scholar?q=Similarity+Estimation+Techniques+from+Rounding+Algorithms
15. A Quantized Johnson-Lindenstrauss Lemma: The Finding of Buffon's Needle — Laurent Jacques, 2015
https://scholar.google.com/scholar?q=A+Quantized+Johnson-Lindenstrauss+Lemma%3A+The+Finding+of+Buffon%27s+Needle
16. Quantized Random Projections and Non-Linear Estimation of Cosine Similarity — Ping Li, Michael Mitzenmacher, Martin Slawski, 2016
https://scholar.google.com/scholar?q=Quantized+Random+Projections+and+Non-Linear+Estimation+of+Cosine+Similarity
17. QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead — Amir Zandieh, Majid Daliri, Insu Han, 2025
https://scholar.google.com/scholar?q=QJL%3A+1-Bit+Quantized+JL+Transform+for+KV+Cache+Quantization+with+Zero+Overhead
18. PolarQuant: Quantizing KV Caches with Polar Transformation — Iman Han, Praveen Kacham, Amin Karbasi, Vahab Mirrokni, and Amir Zandieh, 2025
https://scholar.google.com/scholar?q=PolarQuant%3A+Quantizing+KV+Caches+with+Polar+Transformation
19. Practical and Asymptotically Optimal Quantization of High-Dimensional Vectors in Euclidean Space for Approximate Nearest Neighbor Search — Jianqiu Gao, Yuxuan Gou, Yifan Xu, Yifan Yang, Cheng Long, and Raymond Chi-Wing Wong, 2024
https://scholar.google.com/scholar?q=Practical+and+Asymptotically+Optimal+Quantization+of+High-Dimensional+Vectors+in+Euclidean+Space+for+Approximate+Nearest+Neighbor+Search
20. Accelerating Large-Scale Inference with Anisotropic Vector Quantization — Ruoming Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar, 2020
https://scholar.google.com/scholar?q=Accelerating+Large-Scale+Inference+with+Anisotropic+Vector+Quantization
21. KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache — Zihao Liu, Jun Yuan, Haotian Jin, Sheng Zhong, Ziqi Xu, Vladimir Braverman, Beidi Chen, and Xia Hu, 2024
https://scholar.google.com/scholar?q=KIVI%3A+A+Tuning-Free+Asymmetric+2bit+Quantization+for+KV+Cache
22. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — approx. recent RAG / LLM systems authors, 2025
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
23. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — approx. recent LLM systems authors, 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
24. Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation — approx. recent RAG systems authors, 2025
https://scholar.google.com/scholar?q=Cache-Craft%3A+Managing+Chunk-Caches+for+Efficient+Retrieval-Augmented+Generation
25. Weighted Minwise Hashing Beats Linear Sketching for Inner Product Estimation — approx. sketching / similarity estimation authors, recent
https://scholar.google.com/scholar?q=Weighted+Minwise+Hashing+Beats+Linear+Sketching+for+Inner+Product+Estimation
26. AI Post Transformers: Memory Traffic Saturation in Transformer Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-20-memory-traffic-saturation-in-transformer-cd4961.mp3
27. AI Post Transformers: Quantizing Diffusion LLMs: A Systematic Study — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/quantizing-diffusion-llms-a-systematic-study/
28. AI Post Transformers: Limitations of Embedding-Based Retrieval — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/limitations-of-embedding-based-retrieval/