AI Post Transformers Visual Companion

Stochastic KV Routing for Cache Sharing

Depth-wise KV reuse turns cache memory from a rigid per-layer cost into a serving-time control knob. This page shows the geometry: which layers keep private caches, which layers borrow earlier memory, and where the quality-vs-memory trade begins to bend.

arXiv 2604.22782 Topic Depth-wise cache sharing Bridge GQA → cross-layer reuse Transcript IDs 2604.22782

Paper Focus

Train with stochastic cross-layer routing, then deploy with deterministic retention masks such as every-2nd or every-3rd cache layer. The target is long-context autoregressive decoding under tight HBM budgets.

From Private Cabinets to Shared Depth Memory

Baseline decoding stores separate KV state at every layer. Stochastic KV Routing trains the model to survive partial cache retention, so serving can pick a memory budget without retraining.

Decoder Memory Topology

retained self-cache routed layer borrowing earlier cache HBM pressure

Cache Cost Snapshot

Mock serving profile for a 32-layer decoder, 8k context, batch-heavy decode. Memory falls roughly with retained cache layers, while quality decays more gently in moderate regimes.

Interactive Routing Lab

Slide the retention rate to redraw which layers keep local cache and which layers route backward. Hover cells to inspect routing ownership across depth and sequence positions.

Retention 50%

Depth × Token Cache Heatmap

Each row is a layer, each column a token block. Warmer cells mean heavier dependence on a borrowed cache source rather than a local one.

local / low reuse moderate reuse heavy routed reuse

Memory Saved vs Quality Lost

These curves are illustrative, but they encode the paper’s practical story: the middle band is where routing looks like engineering, while 25% retention starts to expose sharper quality cracks.

Retention Curves

stochastic routing trained post-hoc deterministic sharing

Serving Envelope

Mock normalized values for memory, batch headroom, and robustness under deterministic masks. The aim is not perfect invariance, but graceful degradation across a family of layouts.

Where This Sits in the KV Design Landscape

Depth-wise sharing is orthogonal to time-axis eviction and complementary to within-layer head sharing like GQA. The heatmap below compares method families by training adaptation, serving flexibility, and cache-axis target.

Method Matrix

Axis of Attack

Temporal eviction squeezes sequence positions. Depth-wise routing squeezes layer duplication. GQA squeezes KV head count inside a layer.

References