Two copies of the same transformer architecture, different fine-tuned weights, one KV cache shipped across a network. This page visualizes why that naive transfer breaks — and how REUSE + PATCH rebuild a usable cache without a full recompute.
Prefill is compute-bound; decode is memory-bandwidth-bound. Splitting them across machines fixes the compute/memory mismatch — and creates a network problem: how do you get the KV cache from the prefill machine to the decode machine cheaply?
A KV cache built under one set of weights, read by a model with different weights, injects a small per-layer mismatch. Residual connections carry that error forward — each layer's amplification factor sits at or above 1x, so the error compounds multiplicatively across depth.
REUSE alone lets error snowball toward the output layers. PATCH layers act as checkpoints — they reset the error term instead of letting it compound, so the curve keeps getting cut back down.
Paired producer/consumer KV traces on the same prefix are stacked into a joint matrix. SVD factorization finds a shared low-rank code; a ridge-regression encoder (producer) and decoder (consumer) are fit per layer around that code.
A sparse set of transition layers use PATCH instead of REUSE — the producer ships a compressed pre-attention hidden state, and a small MLP aligner reconstructs native K/V on the consumer side. Layer selection: restore-one sensitivity pass, then greedy budgeted search over the candidate pool (Appendix A).
Raw KV transfer collapses under semantic drift; 4-bit quantization stacks its own error on top. SCD lands within a few points of full consumer-side recompute (Oracle).
SCD's continuous quality/latency knob beats DroidSpeak's binary reuse-or-recompute choice per layer.
Key rank compresses harder than Value rank without hurting F1 — Values appear to carry more of the fine-grained semantic signal.
LoRA-adapter fleets and speculative-decoding draft-verifier pairs justify the whole framework in the introduction — then neither appears in the experiments. Both tested pairs are base-model-to-full-fine-tune.
9.74 hours per producer/consumer pair, dominated by Patch aligner training. The paper's amortization pitch assumes one setup serves many consumers — but a fleet of 20 LoRA specialists means 20 separate ten-hour calibration runs.