Shared round summaries are reused once, then represented as a master cache plus sparse per-agent differences.
Plus up to 1.9x faster prefill than per-request PIC and up to 2.7x more concurrent agents than prefix caching under a latency target.
Use the step buttons to animate the round. The same shared bundle lands at different positions because each agent carries a different private history.
Prefix reuse only sees overlap from token zero. TokenDance sees the entire synchronized round and amortizes matching once.
Hover cells to inspect how much of each agent’s logical KV view is shared. Toggle storage mode to compare duplicate copies against diff-aware storage.
When pairwise block similarity stays in the 90% range, storing one master cache and thin deltas is much cheaper than storing N dense siblings.
Switch workloads to compare realistic mock measurements inspired by the episode: memory, concurrent agents under SLO, and prefill speedup.
The paper’s value is not “generic throughput magic.” It targets persistent KV saturation in synchronized multi-agent rounds.
Two interactive views show sensitivity to token alignment and synchronization. TokenDance is strongest when shared text is exact and the round is available together.
As formatting noise rises or asynchronous stragglers dominate, the collective object becomes less coherent and per-request or subgroup policies look more attractive.