A KV cache is shaped by the exact depth, width, and head count of the model that built it — so handing one model's cache to a different architecture requires a learned translation, not a copy. This page visualizes why naive transfer fails, how Mixture-of-Translators routes tokens through specialized translator experts, and where the paper's honest partial-success numbers land.
Every attention layer computes a key and value vector per token; stacked across layers, that's the KV cache — a model's compressed memory of everything it has read. That cache is baked into the exact shape of the model that produced it, so handing it to a different architecture needs a five-stage translation pipeline, not a copy-paste.
Instead of one universal projection (what Cache-to-Cache, KVComm, LSC, and Interlat all rely on), MoT trains several translator experts and a per-token gate that routes each token to its own top-K blend. A second loss term then corrects drift the translation leaves behind in the target model's own layers.
Sliding the translation window from layer 0 to layer 6 on a homogeneous GPT-2→GPT-2 setup produces a genuine U-shaped validation loss curve, backed by a Lipschitz argument: early injection lets error compound through every remaining layer; late injection leaves too few layers to correct it.
MoT nearly matches its own homogeneous performance even when the source jumps from 0.5B to 7B parameters — but every heterogeneous row compresses a bigger source into the same 0.5B target, not the harder "small model borrows big model's understanding" direction.
Translators train on 128-token windows yet get evaluated up to 24K tokens — nearly two orders of magnitude beyond training length. MoT's F1 climbs rather than collapses, while offloading non-hub caches keeps multi-agent memory flat as the fleet grows.