Most memory sits in long-prefix KV tensors. The proposed split stores one shared backbone cache and small rank-space traces per agent, or fully shared rank traces in shared-A mode.
The key storage trick is to cache XA instead of full-width ΔY=(XA)B. Hover cells to inspect dimensions and which parts are shared.
Mock layer×token similarity illustrates the paper’s premise: on common prefix tokens, backbone-derived states remain highly aligned while adapter deltas explain most deviation.
Instead of materializing full LoRA-expanded K/V in memory, the kernel reconstructs adapter effects inside the attention tile. Step through the fused dataflow.
Mock result curves reflect the episode’s framing: biggest gains appear with long context and many agents. Toggle which metric is emphasized.
An attribution-style stacked bar: conceptual cache sharing cuts residency, while the fused kernel protects runtime by avoiding full reconstruction traffic.
Shared-A often sits on a better frontier in the discussed results: more sharing, less memory, and small performance delta relative to non-shared serving.
Use the sliders to vary context length, number of agents, and shared-prefix ratio. The visualization recomputes estimated cache volume for three serving styles.
Compact source list from the episode and transcript. Additional arXiv IDs found in transcript: 2602.01053.