Two Directions, One Mapper
Cache transfer runs both ways within a matched-KV family pair. Toggle direction to see which cost each one skips.
Variance Explained (Qwen3 14B → 32B)
One source layer vs. stacking multiple source layers via top-k selection.
What Makes A Pair Work
Three-Part Closed-Form Mapper
No backprop. Fit once on ~500 calibration sequences, then applied per head at inference time.
Top-K Source Layer Selection
For each target layer, the mapper doesn't use one fixed source layer — it picks the most predictive ones and concatenates them. Hover a cell to see predictive weight.
Accuracy Retention vs. Standalone Model
Six matched-KV pairs. Four cluster high; two — both involving Ministral 14B — collapse.
GSM8K Is The Outlier — Even In A Clean Pair
Qwen3 8B → 32B looks pristine on single-pass scoring benchmarks, then drops hard on multi-step chain-of-thought generation.
Cross-Layer Selection (k)
Dropping k from 8 to 1 is the single biggest hit to fit quality in the whole mapper.
RoPE Handling At Inference
Disabling RoPE re-rotation is brutal — but only where it matters.
Rescuing Failed Pairs: Ridge vs. Small MLP
Same calibration data. Swapping the mapper's functional form from ridge regression to a small MLP rescues both collapsed Ministral-14B pairs.
Method Landscape
| Method | Gradient Training | Mechanism | Notes |
|---|---|---|---|
| This paper | No | Closed-form ridge regression, per head | Restricted to matched-KV pairs within a family |
| Cache-to-cache (C2C) | Yes | Trained neural fuser per model pair | Not limited to matched-KV shapes |
| LatentAlign | Yes | Learned adapters into shared latent space | Cross-model, needs training run per pair |
| DroidSpeak | No | KV sharing across fine-tuned variants | Easier problem — identical architecture, same base weights |
Speedup vs. Re-Prefill
What Actually Predicts Success?
Fit quality (R²) barely correlates with downstream accuracy across pairs. Attention-output cosine does.
Directional Asymmetry, Same Fit Quality
Llama 3.1 8B↔70B fits with identical R² = 0.84 in both directions — but retention diverges wildly.
References
- 1.Heo, Shafipour, Zhao, Golub, Kamani, Borkar, Chandran, Zardoshti, Rouhani — Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse, 2026. arXiv:2608.03893
- 2.Fu, Min, Zhang, Yan, Dai, Ouyang, Wang — Cache-to-cache: Direct semantic communication between large language models (C2C), ICLR 2026. Scholar
- 3.Dery, Yahav, Prior, Feng, Shen, Szlam — Latent space communication via K-V cache alignment (LatentAlign), 2026. Scholar
- 4.Liu et al. — DroidSpeak: KV cache sharing across fine-tuned model variants, NSDI 2026. Scholar
- 5.Huh, Cheung, Wang, Isola — The Platonic Representation Hypothesis, ICML 2024. Scholar
- 6.Blakeman et al. — Nvidia Nemotron 3: Efficient and open intelligence, 2025. Scholar