Semantic Cache Distillation: Solving Semantic Drift in KV Cache Transfer

arXiv:2606.07684 Qianli Ma, Zhiqing Tang, Hanshuai Cui, Zhi Yao, Weijia Jia — 2026 AI Post Transformers

Two copies of the same transformer architecture, different fine-tuned weights, one KV cache shipped across a network. This page visualizes why that naive transfer breaks — and how REUSE + PATCH rebuild a usable cache without a full recompute.

Prefill/Decode Disaggregation, and the Wire in Between

Prefill is compute-bound; decode is memory-bandwidth-bound. Splitting them across machines fixes the compute/memory mismatch — and creates a network problem: how do you get the KV cache from the prefill machine to the decode machine cheaply?

Semantic Drift: A Small Error, Compounding Through Depth

A KV cache built under one set of weights, read by a model with different weights, injects a small per-layer mismatch. Residual connections carry that error forward — each layer's amplification factor sits at or above 1x, so the error compounds multiplicatively across depth.

low drift moderate drift severe drift PATCH layer (error reset)

Cumulative Error by Layer

REUSE alone lets error snowball toward the output layers. PATCH layers act as checkpoints — they reset the error term instead of letting it compound, so the curve keeps getting cut back down.

REUSE: Shared Low-Rank Latent Code via SVD

Paired producer/consumer KV traces on the same prefix are stacked into a joint matrix. SVD factorization finds a shared low-rank code; a ridge-regression encoder (producer) and decoder (consumer) are fit per layer around that code.

Which Layers Get PATCH?

A sparse set of transition layers use PATCH instead of REUSE — the producer ships a compressed pre-attention hidden state, and a small MLP aligner reconstructs native K/V on the consumer side. Layer selection: restore-one sensitivity pass, then greedy budgeted search over the candidate pool (Appendix A).

REUSE (linear projection) PATCH (MLP aligner, error reset)

Quality Recovery vs. Oracle

Raw KV transfer collapses under semantic drift; 4-bit quantization stacks its own error on top. SCD lands within a few points of full consumer-side recompute (Oracle).

Time-to-First-Token Speedup

SCD's continuous quality/latency knob beats DroidSpeak's binary reuse-or-recompute choice per layer.

Key vs. Value Rank Ablation

Key rank compresses harder than Value rank without hurting F1 — Values appear to carry more of the fine-grained semantic signal.

The Motivating Scenarios Were Never Tested

LoRA-adapter fleets and speculative-decoding draft-verifier pairs justify the whole framework in the introduction — then neither appears in the experiments. Both tested pairs are base-model-to-full-fine-tune.

Calibration Is Per-Consumer, Not Per-Producer

9.74 hours per producer/consumer pair, dominated by Patch aligner training. The paper's amortization pitch assumes one setup serves many consumers — but a fleet of 20 LoRA specialists means 20 separate ten-hour calibration runs.

9.74h
total calibration time / pair
1.58GB
PATCH artifacts (vs 8.4MB REUSE)
397.7M
extra params — 5.49% of 7.24B consumer

References

  1. Ma, Q., Tang, Z., Cui, H., Yao, Z., Jia, W. — Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching, 2026 — arxiv.org/abs/2606.07684source paper
  2. Sheng, Y., Cao, S., Li, D., Hooper, C., Lee, N., Yang, S., Chou, C., Zhu, B., Zheng, L., Keutzer, K., Gonzalez, J.E., Stoica, I. — S-LoRA: Serving Thousands of Concurrent LoRA Adapters, 2024 — scholar link
  3. Liu, Y., Li, H., Cheng, Y., Ray, S., Huang, Y., Zhang, Q., Du, K., Yao, J., Lu, S., Ananthanarayanan, G., Maire, M., Hoffmann, H., Holtzman, A., Jiang, J. — CacheGen: KV Cache Compression and Streaming for Fast LLM Serving, 2024 — scholar link
  4. Yao, J., Li, H., Liu, Y., Ray, S., Cheng, Y., Zhang, Q., Du, K., Lu, S., Jiang, J. — CacheBlend: Fast LLM Serving with Cached Knowledge Fusion, 2024 — scholar link
  5. Leviathan, Y., Kalman, M., Matias, Y. — Fast Inference from Transformers via Speculative Decoding, 2023 — scholar link
  6. Liu, Y., Huang, Y., Yao, J., Feng, S., Gu, Z., Du, K., Li, H., Cheng, Y., Jiang, J., Lu, S., et al. — DroidSpeak: KV Cache Sharing for Cross-LLM Communication, 2024 — scholar link
  7. Fu, T., Min, Z., Zhang, H., Yan, J., Dai, G., Ouyang, W., Wang, Y. — Cache-to-Cache: Direct Semantic Communication Between LLMs, 2026 — scholar link
  8. Sheng, Y., Cao, S., Li, D., Hooper, C., Lee, N., Yang, S., Chou, C., Zhu, B., Zheng, L., Keutzer, K., et al. — S-LoRA: Scalable Serving of Thousands of LoRA Adapters, 2024 — scholar link
  9. Li, Y., Wei, F., Zhang, C., Zhang, H. — EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty, 2024 — scholar link
  10. Qin, R., Li, Z., He, W., Zhang, M., Wu, Y., Zheng, W., Xu, X. — Mooncake: A KVCache-Centric Disaggregated Architecture for LLM Serving, 2024 — scholar link