Mixture-of-Translators: Sharing KV Caches Across Different LLMs

Jin-woo Lee, Minkyung Song, Junghyun Oh, Seunghoon Han, Soyoung Park, Gwangseon Jang, Sungsu Lim — Chungnam National University & KISTI · posted 2026-07-31
arXiv 2607.28979 Topic KV Cache Translation Method Mixture-of-Experts Routing Venue AI Post Transformers

A KV cache is shaped by the exact depth, width, and head count of the model that built it — so handing one model's cache to a different architecture requires a learned translation, not a copy. This page visualizes why naive transfer fails, how Mixture-of-Translators routes tokens through specialized translator experts, and where the paper's honest partial-success numbers land.

The Handoff Problem

Every attention layer computes a key and value vector per token; stacked across layers, that's the KV cache — a model's compressed memory of everything it has read. That cache is baked into the exact shape of the model that produced it, so handing it to a different architecture needs a five-stage translation pipeline, not a copy-paste.

Five-Stage Translation Pipeline

Context is prefilled once on the source model, translated by MoT into the target's shape, replayed through the target's own layers under Context Correction supervision, then the prompt and completion proceed natively.

Same Text, Two Incompatible Cache Shapes

Hover a cell to inspect its (layer, head) coordinate and mock activation magnitude. A 7B source model's cache has far more layers and heads than a 0.5B target — copying tensors directly is a dimension mismatch before it's even a semantic one.

Mixture-of-Translators + Context Correction

Instead of one universal projection (what Cache-to-Cache, KVComm, LSC, and Interlat all rely on), MoT trains several translator experts and a per-token gate that routes each token to its own top-K blend. A second loss term then corrects drift the translation leaves behind in the target model's own layers.

Per-Token Expert Routing (click a token)

Each token gets its own top-2 blend of translator experts — unlike a single shared mapping, different tokens at different layer depths can route differently.

Context Correction Loss: Pulling the Replay Back to Native

Divergence between the target's replayed hidden states and what it would have produced natively, tracked layer-by-layer from the injection point. CC loss directly supervises this gap back toward zero.

The U-Shaped Injection Tradeoff

Sliding the translation window from layer 0 to layer 6 on a homogeneous GPT-2→GPT-2 setup produces a genuine U-shaped validation loss curve, backed by a Lipschitz argument: early injection lets error compound through every remaining layer; late injection leaves too few layers to correct it.

Validation Loss vs. Injection Layer

Total Loss Propagation Bound Correction Deficit
Propagation risk (1+δ)^(remaining layers) dominates early injection; correction-deficit coefficient climbs toward 1 near the final layer. Their sum is the U — no injection point is fully safe on its own.

Accuracy & F1 Across Settings

MoT nearly matches its own homogeneous performance even when the source jumps from 0.5B to 7B parameters — but every heterogeneous row compresses a bigger source into the same 0.5B target, not the harder "small model borrows big model's understanding" direction.

Same-Size vs. Heterogeneous Cache Transfer

Average accuracy — BoolQ / PubMedQA / MMLU-Redux (%)
F1 — SQuAD / NewsQA (extractive QA)
Native (52.0% / 0.45 F1), MoT (49.0%→51.0% / 0.42→0.43 F1), and C2C-Project's ≈0.03 F1 collapse are the paper's reported figures. Interlat, KVComm, and LSC bars illustrate the qualitative trends described in the transcript ("partial recovery," "fine on BoolQ, not run heterogeneous," "degrades across the board") — exact values for those three beyond the described trend were not given.

Long-Context Generalization & Multi-Agent Memory

Translators train on 128-token windows yet get evaluated up to 24K tokens — nearly two orders of magnitude beyond training length. MoT's F1 climbs rather than collapses, while offloading non-hub caches keeps multi-agent memory flat as the fleet grows.

F1 vs. Context Budget (Cache-Augmented Generation)

Trained on 64-token context windows; evaluated from 4K to 24K tokens. MoT improves with scale instead of degrading — a result the paper reports without a full mechanistic explanation.

Peak Memory vs. Number of Agents

Retain mode: peak memory grows ~77–86% as agents scale for all three methods. Free mode (MoT only, as reported): a Hub Agent keeps full history while every other agent offloads after its turn — at 10 agents MoT-Free sits near 0.3GB vs. >1GB for Retain, with F1 unchanged.

References