← All episodes CacheBridge: Fixing Cross-Model KV Cache Transfer Failures

CacheBridge: Fixing Cross-Model KV Cache Transfer Failures

Sep 17, 2026
This episode examines CacheBridge, a paper proposing targeted fixes to a training-free method for transferring KV caches between different transformer models in multi-model routing setups. It explains why caches can't simply be handed off — differing residual widths, GQA head counts, and RoPE position encoding make one model's cache unreadable to another — and how a prior affine-mapper approach (FULL-HEADMAPPING) could swing wildly from near-native accuracy on one model pair to catastrophic collapse on another, with no way to predict which. The discussion breaks down the paper's four diagnosed failure causes, spanning head-mixing, mismatched error metrics, layer-count cost scaling, and a GPU implementation bottleneck, then covers the three corresponding repairs: HEAD-LOCAL's narrower one-to-one head mapping, ATTN-REPAIR's attention-aware calibration reweighting, and FUSED-FIT's custom kernel for building the mapper efficiently. Listeners interested in LLM serving infrastructure will find it a concrete look at diagnosing and patching a deployed technique rather than proposing a new architecture from scratch.
Sources:
1. CacheBridge: Efficient Cross-Model KV Cache Transfer — Xingyu Qu, Siyuan Lu, Zhiyu Chen, Sheng Wang, Tao Lin, 2026
http://arxiv.org/abs/2609.00891
2. Cross-model KV cache transfer in LLM families: A closed-form linear mapping for prefill reuse — Heo, T., Shafipour, R., Zhao, R., Golub, M., Kamani, M. M., Borkar, R., Chandran, M. T., Zardoshti, P., Rouhani, B. D., 2026
https://scholar.google.com/scholar?q=Cross-model+KV+cache+transfer+in+LLM+families%3A+A+closed-form+linear+mapping+for+prefill+reuse
3. Cache-to-cache: Direct semantic communication between large language models — Fu, T., Min, Z., Zhang, H., Yan, J., Dai, G., Ouyang, W., Wang, Y., 2026
https://scholar.google.com/scholar?q=Cache-to-cache%3A+Direct+semantic+communication+between+large+language+models
4. Mixture-of-translators: Translating KV caches across heterogeneous large language models — Lee, J.-w., Song, M., Oh, J., Han, S., Park, S., Jang, G., Lim, S., 2026
https://scholar.google.com/scholar?q=Mixture-of-translators%3A+Translating+KV+caches+across+heterogeneous+large+language+models
5. DroidSpeak: KV cache sharing across fine-tuned model variants — Liu, Y., Huang, Y., Yao, J., Feng, S., Gu, Z., Du, K., Li, H., Cheng, Y., Jiang, J., Lu, S., Musuvathi, M., Choukse, E., 2026
https://scholar.google.com/scholar?q=DroidSpeak%3A+KV+cache+sharing+across+fine-tuned+model+variants
6. ICaRus: Identical cache reuse for efficient multi model inference — Woo, S., Kil, J., Kim, H., Kim, M., Kim, J., Seo, A., Lee, S., Jo, M., Ryu, J., Park, B., Kwon, S. J., Lee, D., 2026
https://scholar.google.com/scholar?q=ICaRus%3A+Identical+cache+reuse+for+efficient+multi+model+inference
7. GQA: Training generalized multi-query transformer models from multi-head checkpoints — Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., Sanghai, S., 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+generalized+multi-query+transformer+models+from+multi-head+checkpoints
Interactive Visualization: CacheBridge: Fixing Cross-Model KV Cache Transfer Failures