Interactive Episode Companion

TokenDance for Multi-Agent KV Cache Sharing

A visual walkthrough of why synchronized multi-agent rounds create extreme KV duplication, and how TokenDance reframes reuse from single requests to the whole round as one collective object.
arXiv: 2604.03143 2026 • Bian, Wu, Zhang, Dong, Liang, Zhuo Workload: all-gather agent rounds Core bottleneck: persistent KV memory Transcript IDs found: 2604.03143 only

Claim Shape

round-level reuse > request-level reuse

Shared round summaries are reused once, then represented as a master cache plus sparse per-agent differences.

Headline Outcomes

11-17x KV compression

Plus up to 1.9x faster prefill than per-request PIC and up to 2.7x more concurrent agents than prefix caching under a latency target.

Collective Reuse Pipeline

Use the step buttons to animate the round. The same shared bundle lands at different positions because each agent carries a different private history.

What changes

Prefix reuse only sees overlap from token zero. TokenDance sees the entire synchronized round and amortizes matching once.

Sibling Cache Similarity Heatmap

Hover cells to inspect how much of each agent’s logical KV view is shared. Toggle storage mode to compare duplicate copies against diff-aware storage.

low overlapshared spannear-identical

Storage picture

When pairwise block similarity stays in the 90% range, storing one master cache and thin deltas is much cheaper than storing N dense siblings.

Service Outcomes Under Two Agent Traces

Switch workloads to compare realistic mock measurements inspired by the episode: memory, concurrent agents under SLO, and prefill speedup.

Interpretation

The paper’s value is not “generic throughput magic.” It targets persistent KV saturation in synchronized multi-agent rounds.

Where The Method Holds And Where It Frays

Two interactive views show sensitivity to token alignment and synchronization. TokenDance is strongest when shared text is exact and the round is available together.

Reading the boundary

As formatting noise rises or asynchronous stragglers dominate, the collective object becomes less coherent and per-request or subgroup policies look more attractive.

References

TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
Bian et al., 2026 • arXiv:2604.03143
vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention
Kwon et al., 2023 • arXiv:2309.06180
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Dao et al., 2022 • arXiv:2205.14135
Generative Agents: Interactive Simulacra of Human Behavior
Park et al., 2023 • arXiv:2304.03442