Efficient KV Cache Sharing for Multi-LoRA Agents

A visual companion for the podcast episode on LRAgent: how multi-agent systems can share one large backbone KV cache while keeping only tiny low-rank, role-specific traces.
arXiv: 2602.01053 Topic: KV cache reuse LoRA rank-space caching Flash-LoRA-Attention Shared-A multi-LoRA

Shared trunk, tiny role branches

Most memory sits in long-prefix KV tensors. The proposed split stores one shared backbone cache and small rank-space traces per agent, or fully shared rank traces in shared-A mode.

shared backbone KV agent-specific low-rank cache up-projection / reconstruction reused computation

LoRA cache decomposition

The key storage trick is to cache XA instead of full-width ΔY=(XA)B. Hover cells to inspect dimensions and which parts are shared.

Similarity heatmap across agents

Mock layer×token similarity illustrates the paper’s premise: on common prefix tokens, backbone-derived states remain highly aligned while adapter deltas explain most deviation.

low mid high

Flash-LoRA-Attention reconstruction path

Instead of materializing full LoRA-expanded K/V in memory, the kernel reconstructs adapter effects inside the attention tile. Step through the fused dataflow.

Memory, TTFT, throughput

Mock result curves reflect the episode’s framing: biggest gains appear with long context and many agents. Toggle which metric is emphasized.

Where does the gain come from?

An attribution-style stacked bar: conceptual cache sharing cuts residency, while the fused kernel protects runtime by avoiding full reconstruction traffic.

Accuracy vs efficiency frontier

Shared-A often sits on a better frontier in the discussed results: more sharing, less memory, and small performance delta relative to non-shared serving.

How overlap-dependent is this?

Use the sliders to vary context length, number of agents, and shared-prefix ratio. The visualization recomputes estimated cache volume for three serving styles.

Interactive model
shared_prefix_ratio
80%
context_tokens
32768
num_agents
8
Non-shared KV BaseShared BaseLR-Shared

References

Compact source list from the episode and transcript. Additional arXiv IDs found in transcript: 2602.01053.

LRAgent — Jeon, Ha, Kim (2026)
arXiv:2602.01053
LoRA — Hu et al. (2022)
Scholar link
FlashAttention — Dao et al. (2022)
Scholar link
FlashAttention-2 — Dao (2023)
Scholar link
S-LoRA — Wang et al. (2023)
Scholar link
MiLoRA — Xia et al. (2024)
Scholar link
MELoRA — Tian et al. (2024)
Scholar link
Multi-Head Latent Attention — DeepSeek-AI team (2025)
Scholar link
KV Packet
Scholar link