← All episodes Mixture-of-Translators: Sharing KV Caches Across Different LLMs

Mixture-of-Translators: Sharing KV Caches Across Different LLMs

Sep 17, 2026
This episode examines "Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models" by Jin-woo Lee and six co-authors from Chungnam National University and KISTI, which tackles the problem of transferring one model's KV cache — its layer-by-layer key/value memory of a processed document — to a completely different model architecture without re-running the original text through it. The discussion explains why this is hard: a KV cache is shaped by the specific depth, width, and head count of the model that produced it, so naive copying fails and a learned "cache translation" is required instead. It surveys prior approaches (Cache-to-Cache, KVComm, Latent Space Communication, and Interlat) and their shared weakness — relying on a single universal mapping or shared latent space — before detailing how Mixture-of-Translators borrows the Mixture-of-Experts routing idea to assign different tokens to different specialized translator modules via a per-token gating network, paired with a Context Correction Loss to correct drift in the target model's own layers. Listeners interested in multi-agent LLM pipelines, cache-augmented generation, or reducing redundant prefill computation across heterogeneous model fleets will find the practical motivation and technical tradeoffs compelling, especially since the paper's honest partial-success framing offers more insight than a clean win would.
Sources:
1. Mixture-of-Translators: Sharing KV Caches Across Different LLMs
https://arxiv.org/pdf/2607.28979
2. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, Jeff Dean, 2017
https://scholar.google.com/scholar?q=Outrageously+Large+Neural+Networks%3A+The+Sparsely-Gated+Mixture-of-Experts+Layer
3. Relative Representations Enable Zero-Shot Latent Space Communication — Luca Moschella, Valentino Maiorca, Marco Fumero, Antonio Norelli, Francesco Locatello, Emanuele Rodolà, 2023
https://scholar.google.com/scholar?q=Relative+Representations+Enable+Zero-Shot+Latent+Space+Communication
4. Prompt Cache: Modular Attention Reuse for Low-Latency Inference — In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, Lin Zhong, 2024
https://scholar.google.com/scholar?q=Prompt+Cache%3A+Modular+Attention+Reuse+for+Low-Latency+Inference
5. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
6. Cache-to-cache: Direct semantic communication between large language models — Tianyu Fu, Zihan Min, Hanling Zhang, Jichao Yan, Guohao Dai, Wanli Ouyang, Yu Wang, 2025
https://scholar.google.com/scholar?q=Cache-to-cache%3A+Direct+semantic+communication+between+large+language+models
7. KVComm: Enabling efficient LLM communication through selective KV sharing — Xiangyu Shi, Marco Chiesa, Gerald Q Maguire Jr, Dejan Kostic, 2025
https://scholar.google.com/scholar?q=KVComm%3A+Enabling+efficient+LLM+communication+through+selective+KV+sharing
8. Enabling agents to communicate entirely in latent space (Interlat) — Zhuoyun Du, Runze Wang, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Yu Cheng, Bo Zheng, Wei Chen, Haochao Ying, 2025
https://scholar.google.com/scholar?q=Enabling+agents+to+communicate+entirely+in+latent+space+%28Interlat%29
9. Latent space communication via KV cache alignment (LSC) — Lucio M Dery, Zohar Yahav, Henry Prior, Qixuan Feng, Jiajun Shen, Arthur Szlam, 2026
https://scholar.google.com/scholar?q=Latent+space+communication+via+KV+cache+alignment+%28LSC%29
10. Fast state restoration in LLM serving with HCache — Shiwei Gao, Youmin Chen, Jiwu Shu, 2025
https://scholar.google.com/scholar?q=Fast+state+restoration+in+LLM+serving+with+HCache
Interactive Visualization: Mixture-of-Translators: Sharing KV Caches Across Different LLMs