A visual companion to the episode on segment-level KV reuse: when many agents reread the same plan, critique, or summary, the real bottleneck is prefill. The paper’s promise is shared transformer working state; the friction is that those activations are tied to position, not just meaning.
Prefix caching is stable because token order matches exactly. Multi-agent workflows break that assumption by wrapping the same artifact inside different role prompts and dialogue history.
The same plan appears inside solver, critic, summarizer, and reviser prompts. Reuse is easy if the artifact stays at the front; it becomes fragile once it lands at a different offset.
The cache is not just text memory. It is a position-sensitive internal state. Shift a segment to a new offset, and attention similarity can degrade even when the tokens are identical.
Exact-prefix reuse preserves the diagonal. Shifted spans bend it. Selective recompute restores structure, but eats into the latency win.
These mock curves separate three claims: lower prefill time, higher throughput, and the murkier possibility that shared state changes agent behavior rather than only preserving it.
If the workflows repeatedly circulate exact artifacts, gains can be large without proving semantic equivalence. The chart lets you switch between templated and paraphrased workloads.
The systems contribution lives in the serving stack: detect a reusable span, map its logical hash to already materialized physical KV pages, and keep decoding without copying the whole cache.
The idea is not “semantic understanding of any similar passage.” It is a runtime index plus page-backed aliasing, which is powerful when the same artifacts recur in many wrappers.