CacheBridge: Fixing Cross-Model KV Cache Transfer Failures

Xingyu Qu, Siyuan Lu, Zhiyu Chen, Sheng Wang, Tao Lin · Westlake University / Wuhan University / Amazon · arXiv preprint, Sep 1 2026
arXiv:2609.00891

A prior training-free method for handing a KV cache between two models scored 77.2% natively but collapsed to 44.4% when transplanted — on a different model pair the same method was nearly perfect. CacheBridge diagnoses why, and patches it with three targeted repairs.

Why a cache can't just be handed off

Each transformer's KV cache is architecture-specific: residual stream width, GQA head count, and RoPE-encoded position all differ across models. The receiving model literally cannot read the sender's cache without translation.

Sender-side state Affine mapper (offline, ridge-fit) Receiver-side state

Four diagnosed failure causes

FULL-HEADMAPPING — the prior baseline — fans every target head in from all source heads. The authors trace its instability to four distinct problems.

01

Head-mixing

Full fan-in draws on every source KV head, diluting the signal a target head actually needs.

02

Mismatched error metrics

Coordinate-level R² fit does not predict downstream generation quality.

03

Layer-count cost scaling

Pulling from more layers pushes mapper cost toward the price of just re-prefilling.

04

GPU implementation bottleneck

Scattered per-layer source sets make efficient GPU construction awkward and slow.

The collapse, in numbers

Three repairs, one deployed interface

Same affine, training-free mapper — nothing about the online interface changes. Each repair targets one of the four causes above.

HEAD-LOCAL

One-to-one head mapping instead of all-to-one fan-in. Fixes cause 01, and shrinks cause 03's cost.

ATTN-REPAIR

Reweights calibration by the receiver's real attention sensitivity, not raw coordinate error. Fixes cause 02.

FUSED-FIT

A two-pass GPU kernel that streams scattered layer supports without materializing them. Fixes cause 04.

All-to-one vs. one-to-one head assignment

Every model tested exposes 8 KV heads. FULL-HEADMAPPING lets each target head draw from all 8 source heads per selected layer. HEAD-LOCAL sets a(h) = h: a fixed, architecture-indexed, deterministic one-to-one map.

Rows = target heads, columns = source heads. Colored cell = this source head contributes to this target head's mapper.

Feature width per target head

k = selected layers, Hs = 8 source heads, ds = per-head dim. Full fan-in scales with Hs; HEAD-LOCAL doesn't.

Why narrower is safer, not just cheaper

Full fan-in raises the ridge fit's effective degrees of freedom, retaining cross-head directions only weakly identified by calibration data.

Every mapped layer's cache error propagates through the receiver's downstream Jacobian — an early overfit direction compounds all the way to the output. Restricting support is a stability fix as much as an efficiency one.

8×
feature width reduction
Hs=8
KV heads, every model tested

Test bed

Three GQA-to-GQA directions — Ministral 3B→14B, Ministral 8B→14B, Qwen3 14B→32B — 500 FineWeb-Edu calibration sequences, evaluated on HellaSwag, ARC-Challenge, WinoGrande, and MMLU.

HellaSwag accuracy: FULL-HEADMAPPING vs. CacheBridge

Bars show HellaSwag accuracy where FULL-HEADMAPPING was collapsing. CacheBridge = HEAD-LOCAL + ATTN-REPAIR.

On Qwen3 14B→32B, where the baseline wasn't broken, CacheBridge doesn't regress mean target retention.

Is it really the head correspondence?

Two controls at matched parameter budget on Qwen3 14B→32B: block PCA over all source heads, and cyclic shifts of the head assignment away from identity.

Illustrative bars within the paper's reported range: PCA loses ~14 retention points; every cyclic shift loses 13–15+ points versus the aligned (identity) assignment — despite matched coefficient counts.

References

1CacheBridge: Efficient Cross-Model KV Cache Transfer — Qu, Lu, Chen, Wang, Lin, 2026
2Cross-model KV cache transfer in LLM families: A closed-form linear mapping for prefill reuse — Heo, Shafipour, Zhao, Golub, Kamani, Borkar, Chandran, Zardoshti, Rouhani, 2026 (FULL-HEADMAPPING baseline)
3Cache-to-cache: Direct semantic communication between large language models — Fu, Min, Zhang, Yan, Dai, Ouyang, Wang, 2026
4Mixture-of-translators: Translating KV caches across heterogeneous large language models — Lee, Song, Oh, Han, Park, Jang, Lim, 2026
5DroidSpeak: KV cache sharing across fine-tuned model variants — Liu, Huang, Yao, Feng, Gu, Du, Li, Cheng, Jiang, Lu, Musuvathi, Choukse, 2026
6ICaRus: Identical cache reuse for efficient multi model inference — Woo, Kil, Kim, Kim, Kim, Seo, Lee, Jo, Ryu, Park, Kwon, Lee, 2026
7GQA: Training generalized multi-query transformer models from multi-head checkpoints — Ainslie, Lee-Thorp, de Jong, Zemlyanskiy, Lebrón, Sanghai, 2023