AI Post Transformers · Episode Companion

Cross-Model KV Cache Transfer for Fast LLM Prefill Reuse

arXiv:2608.03893 NVIDIA · 2026 Ridge Regression · No Gradient Training

A single layer of a smaller model's KV cache explains over half the variance in a larger model's keys — with zero gradient training. This page visualizes how NVIDIA's closed-form linear mapper transfers KV cache across differently-sized models in the same family, where it wins, and where it quietly breaks.

Two Directions, One Mapper

Cache transfer runs both ways within a matched-KV family pair. Toggle direction to see which cost each one skips.

Variance Explained (Qwen3 14B → 32B)

One source layer vs. stacking multiple source layers via top-k selection.

What Makes A Pair Work

56%
key variance, 1 layer
79%
key variance, stacked
4/6
pairs retain 73–98%
2.7–25×
faster than re-prefill
Precondition: requires a "matched-KV pair" — identical KV head count and per-head dimension. Same family, different depth is fine; different tensor shapes at the head level are not.

Three-Part Closed-Form Mapper

No backprop. Fit once on ~500 calibration sequences, then applied per head at inference time.

Top-K Source Layer Selection

For each target layer, the mapper doesn't use one fixed source layer — it picks the most predictive ones and concatenates them. Hover a cell to see predictive weight.

low contribution medium high contribution

Accuracy Retention vs. Standalone Model

Six matched-KV pairs. Four cluster high; two — both involving Ministral 14B — collapse.

Tier 1 (≥73%) Collapsed (<50%)

GSM8K Is The Outlier — Even In A Clean Pair

Qwen3 8B → 32B looks pristine on single-pass scoring benchmarks, then drops hard on multi-step chain-of-thought generation.

Why: log-likelihood benchmarks score one forward pass. GSM8K generates token-by-token off the mapped cache, so small errors compound across the chain before they show up in the answer.

Cross-Layer Selection (k)

Dropping k from 8 to 1 is the single biggest hit to fit quality in the whole mapper.

RoPE Handling At Inference

Disabling RoPE re-rotation is brutal — but only where it matters.

Rescuing Failed Pairs: Ridge vs. Small MLP

Same calibration data. Swapping the mapper's functional form from ridge regression to a small MLP rescues both collapsed Ministral-14B pairs.

Ridge regression Small MLP

Method Landscape

MethodGradient TrainingMechanismNotes
This paperNoClosed-form ridge regression, per headRestricted to matched-KV pairs within a family
Cache-to-cache (C2C)YesTrained neural fuser per model pairNot limited to matched-KV shapes
LatentAlignYesLearned adapters into shared latent spaceCross-model, needs training run per pair
DroidSpeakNoKV sharing across fine-tuned variantsEasier problem — identical architecture, same base weights
The paper never directly benchmarks accuracy against C2C or LatentAlign — it argues from training cost, not head-to-head quality.

Speedup vs. Re-Prefill

What Actually Predicts Success?

Fit quality (R²) barely correlates with downstream accuracy across pairs. Attention-output cosine does.

Directional Asymmetry, Same Fit Quality

Llama 3.1 8B↔70B fits with identical R² = 0.84 in both directions — but retention diverges wildly.

References