AI Post Transformers • Interactive Visualization

ForkKV for Multi-LoRA Agent Serving

A visual tour of copy-on-write KV caches for branching LoRA agents: where ordinary prefix reuse fails, where residual caches help, and why the real prize is packing more long-context agent branches into fixed GPU memory.

arXiv 2604.06370 Paper ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache Theme KV cache as memory management Mode SVG-only, interactive

One Prefix, Many Agents, One Big Shared Cache

ForkKV treats shared context like a parent process. Each child agent inherits the large base KV state and writes only a small LoRA-specific residual branch.

Shared base KV Agent residual KV Branch-specific decode

Unix Analogy

Parent prompt state is inherited first. Private memory appears only where an agent diverges.

LoRA Twist

Adapters perturb activations, so identical text no longer guarantees identical KV tensors.

Operational Win

Smaller per-branch cache means higher concurrency before GPU memory becomes the hard wall.

Where Ordinary Prefix Caching Stops Working

Same tokens, different adapters, different activations. The first layer admits a clean base-plus-residual split; later layers drift as hidden states diverge.

Interpretation
Layer 1 exact
Shared input activations make the base/residual decomposition exact at the first layer.
Hover signal
ΔKV
Cells encode adapter-induced divergence across layers and token positions.
Mock workload
4 agents
Long shared repo context, then branch-specific tool use and code edits.

Memory Pressure Turns into Throughput Headroom

These charts use realistic mock values shaped by the paper’s reported story: large memory savings, up to roughly 3.0× throughput in favorable branch-heavy settings, and a modest quality tradeoff.

Step-by-Step: DualRadixTree + ResidualAttention

ForkKV needs two things at once: a branching index that finds inherited prefixes fast, and a kernel that reconstructs attention from shared base plus residuals without materializing a giant duplicate KV buffer.

References

Compact links for the main lineage around ForkKV, LoRA serving, and KV-cache-centric systems.