Depth-wise KV reuse turns cache memory from a rigid per-layer cost into a serving-time control knob. This page shows the geometry: which layers keep private caches, which layers borrow earlier memory, and where the quality-vs-memory trade begins to bend.
Train with stochastic cross-layer routing, then deploy with deterministic retention masks such as every-2nd or every-3rd cache layer. The target is long-context autoregressive decoding under tight HBM budgets.
Baseline decoding stores separate KV state at every layer. Stochastic KV Routing trains the model to survive partial cache retention, so serving can pick a memory budget without retraining.
Mock serving profile for a 32-layer decoder, 8k context, batch-heavy decode. Memory falls roughly with retained cache layers, while quality decays more gently in moderate regimes.
Slide the retention rate to redraw which layers keep local cache and which layers route backward. Hover cells to inspect routing ownership across depth and sequence positions.
Each row is a layer, each column a token block. Warmer cells mean heavier dependence on a borrowed cache source rather than a local one.
These curves are illustrative, but they encode the paper’s practical story: the middle band is where routing looks like engineering, while 25% retention starts to expose sharper quality cracks.
Mock normalized values for memory, batch headroom, and robustness under deterministic masks. The aim is not perfect invariance, but graceful degradation across a family of layouts.
Depth-wise sharing is orthogonal to time-axis eviction and complementary to within-layer head sharing like GQA. The heatmap below compares method families by training adaptation, serving flexibility, and cache-axis target.
Temporal eviction squeezes sequence positions. Depth-wise routing squeezes layer duplication. GQA squeezes KV head count inside a layer.