One Prefix, Many Agents, One Big Shared Cache
ForkKV treats shared context like a parent process. Each child agent inherits the large base KV state and writes only a small LoRA-specific residual branch.
Unix Analogy
Parent prompt state is inherited first. Private memory appears only where an agent diverges.
LoRA Twist
Adapters perturb activations, so identical text no longer guarantees identical KV tensors.
Operational Win
Smaller per-branch cache means higher concurrency before GPU memory becomes the hard wall.
Where Ordinary Prefix Caching Stops Working
Same tokens, different adapters, different activations. The first layer admits a clean base-plus-residual split; later layers drift as hidden states diverge.
Memory Pressure Turns into Throughput Headroom
These charts use realistic mock values shaped by the paper’s reported story: large memory savings, up to roughly 3.0× throughput in favorable branch-heavy settings, and a modest quality tradeoff.
Step-by-Step: DualRadixTree + ResidualAttention
ForkKV needs two things at once: a branching index that finds inherited prefixes fast, and a kernel that reconstructs attention from shared base plus residuals without materializing a giant duplicate KV buffer.
References
Compact links for the main lineage around ForkKV, LoRA serving, and KV-cache-centric systems.