AI Post Transformers ICLR 2026 double-blind Systems + Transformer internals

Breaking the Prefix Barrier with Shared KV Cache

A visual companion to the episode on segment-level KV reuse: when many agents reread the same plan, critique, or summary, the real bottleneck is prefill. The paper’s promise is shared transformer working state; the friction is that those activations are tied to position, not just meaning.

5x
agents rereading the same artifact in the mock workflow
72%
synthetic prefill cost share before reuse
3
reuse modes shown here: exact prefix, shifted span, selective recompute
40%
recompute ratio used in the failure-repair illustration

From Exact Prefixes to Shared Segments

Prefix caching is stable because token order matches exactly. Multi-agent workflows break that assumption by wrapping the same artifact inside different role prompts and dialogue history.

reused KV pages fresh wrapper tokens prefill recompute

Reading Pattern

The same plan appears inside solver, critic, summarizer, and reviser prompts. Reuse is easy if the artifact stays at the front; it becomes fragile once it lands at a different offset.

RoPE Misalignment Heatmap

The cache is not just text memory. It is a position-sensitive internal state. Shift a segment to a new offset, and attention similarity can degrade even when the tokens are identical.

low match medium high stress / repair need

Interpretation

Exact-prefix reuse preserves the diagonal. Shifted spans bend it. Selective recompute restores structure, but eats into the latency win.

What the Speedup Is Buying

These mock curves separate three claims: lower prefill time, higher throughput, and the murkier possibility that shared state changes agent behavior rather than only preserving it.

throughput TTFT / prefill task score

Benchmark Lens

If the workflows repeatedly circulate exact artifacts, gains can be large without proving semantic equivalence. The chart lets you switch between templated and paraphrased workloads.

Memory Table and Page Aliasing

The systems contribution lives in the serving stack: detect a reusable span, map its logical hash to already materialized physical KV pages, and keep decoding without copying the whole cache.

Why This Matters

The idea is not “semantic understanding of any similar passage.” It is a runtime index plus page-backed aliasing, which is powerful when the same artifacts recur in many wrappers.

References

Towards a Collaborative Memory for Agentic Workflow: Breaking the Prefix Barrier with Segment-Level KV Cache Sharing OpenReview / ICLR 2026 double-blind openreview.net/forum?id=kgzBkyqg6Z
Efficient Memory Management for Large Language Model Serving with PagedAttention Kwon et al., 2023 Scholar search
SGLang: Efficient Execution of Structured Language Model Programs Zheng et al., 2024 Scholar search
CacheBlend / KVCOMM / EPIC / KVFlow / DroidSpeak / TokenDance 2024–2026 adjacent work on non-prefix reuse, alignment, and multi-agent cache sharing Start with CacheBlend
AI Post Transformers: Efficient KV Cache Sharing for Multi-LoRA Agents Related episode, 2026 Listen
AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing Related episode, 2026 Listen