A prior training-free method for handing a KV cache between two models scored 77.2% natively but collapsed to 44.4% when transplanted — on a different model pair the same method was nearly perfect. CacheBridge diagnoses why, and patches it with three targeted repairs.
Each transformer's KV cache is architecture-specific: residual stream width, GQA head count, and RoPE-encoded position all differ across models. The receiving model literally cannot read the sender's cache without translation.
FULL-HEADMAPPING — the prior baseline — fans every target head in from all source heads. The authors trace its instability to four distinct problems.
Full fan-in draws on every source KV head, diluting the signal a target head actually needs.
Coordinate-level R² fit does not predict downstream generation quality.
Pulling from more layers pushes mapper cost toward the price of just re-prefilling.
Scattered per-layer source sets make efficient GPU construction awkward and slow.
Same affine, training-free mapper — nothing about the online interface changes. Each repair targets one of the four causes above.
One-to-one head mapping instead of all-to-one fan-in. Fixes cause 01, and shrinks cause 03's cost.
Reweights calibration by the receiver's real attention sensitivity, not raw coordinate error. Fixes cause 02.
A two-pass GPU kernel that streams scattered layer supports without materializing them. Fixes cause 04.
Every model tested exposes 8 KV heads. FULL-HEADMAPPING lets each target head draw from all 8 source heads per selected layer. HEAD-LOCAL sets a(h) = h: a fixed, architecture-indexed, deterministic one-to-one map.
Rows = target heads, columns = source heads. Colored cell = this source head contributes to this target head's mapper.
k = selected layers, Hs = 8 source heads, ds = per-head dim. Full fan-in scales with Hs; HEAD-LOCAL doesn't.
Full fan-in raises the ridge fit's effective degrees of freedom, retaining cross-head directions only weakly identified by calibration data.
Every mapped layer's cache error propagates through the receiver's downstream Jacobian — an early overfit direction compounds all the way to the output. Restricting support is a stability fix as much as an efficiency one.
Three GQA-to-GQA directions — Ministral 3B→14B, Ministral 8B→14B, Qwen3 14B→32B — 500 FineWeb-Edu calibration sequences, evaluated on HellaSwag, ARC-Challenge, WinoGrande, and MMLU.
Bars show HellaSwag accuracy where FULL-HEADMAPPING was collapsing. CacheBridge = HEAD-LOCAL + ATTN-REPAIR.
On Qwen3 14B→32B, where the baseline wasn't broken, CacheBridge doesn't regress mean target retention.
Two controls at matched parameter budget on Qwen3 14B→32B: block PCA over all source heads, and cyclic shifts of the head assignment away from identity.
Illustrative bars within the paper's reported range: PCA loses ~14 retention points; every cyclic shift loses 13–15+ points versus the aligned (identity) assignment — despite matched coefficient counts.
| 1 | CacheBridge: Efficient Cross-Model KV Cache Transfer — Qu, Lu, Chen, Wang, Lin, 2026 |
| 2 | Cross-model KV cache transfer in LLM families: A closed-form linear mapping for prefill reuse — Heo, Shafipour, Zhao, Golub, Kamani, Borkar, Chandran, Zardoshti, Rouhani, 2026 (FULL-HEADMAPPING baseline) |
| 3 | Cache-to-cache: Direct semantic communication between large language models — Fu, Min, Zhang, Yan, Dai, Ouyang, Wang, 2026 |
| 4 | Mixture-of-translators: Translating KV caches across heterogeneous large language models — Lee, Song, Oh, Han, Park, Jang, Lim, 2026 |
| 5 | DroidSpeak: KV cache sharing across fine-tuned model variants — Liu, Huang, Yao, Feng, Gu, Du, Li, Cheng, Jiang, Lu, Musuvathi, Choukse, 2026 |
| 6 | ICaRus: Identical cache reuse for efficient multi model inference — Woo, Kil, Kim, Kim, Kim, Seo, Lee, Jo, Ryu, Park, Kwon, Lee, 2026 |
| 7 | GQA: Training generalized multi-query transformer models from multi-head checkpoints — Ainslie, Lee-Thorp, de Jong, Zemlyanskiy, Lebrón, Sanghai, 2023 |