Bidaw: Computation-Storage Aware KV Caching for LLMs

Dramatically improving long, multi-turn AI chatbot performance through bidirectional awareness between GPU scheduler and storage system.

The 93% Computation Waste Problem

Interactive LLM conversations require loading key-value tensors from all previous turns. With an average of 22 rounds per conversation, 93.1% of computation is redundant—recalculating previously computed KVs.

93.1%
Redundant Computation
3.8Ă—
Latency Penalty (Baseline)
50%
Throughput Loss (Baseline)
38s
Median Think Time

Mutual Unawareness: The Root Cause

The GPU scheduler has no idea where KVs are stored (DRAM vs SSD), while the storage system doesn't know which conversations will heat up next. This coordination failure causes massive inefficiency.

Storage Hierarchy Economics

GPU memory is too small. Host DRAM costs $5-10/GB. SSDs cost $0.10-0.20/GB but are 100-1000Ă— slower for random reads. Two-tier caching is essential, but coordination is the key.

Storage-Aware Dual-Queue Scheduling

Bidaw introduces two queues: a ready queue for requests with KVs in DRAM, and a preparing queue for requests loading from SSD. Ready requests never block on I/O.

Preparing Queue Prioritization

Within the preparing queue, requests are reordered by estimated I/O time (primarily KV size). Smaller KVs load faster from SSD and get priority. Waiting time prevents starvation.

LLM-Guided Eviction Policy

Traditional LRU fails because temporal locality is terrible—inter-arrival times are dominated by human think time. Bidaw predicts next access time by analyzing the model's response.

Prediction Features

Weighted Reuse Distance Calculation

Bidaw combines answer length (longer = more reading time), question presence (prompts quicker reply), and sentiment (negative = slower re-engagement) into a weighted reuse distance. This is mapped to hit probability via a ghost cache.

Performance Gains

Bidaw achieves up to 3.58Ă— latency reduction and 1.83Ă— throughput improvement, approaching the theoretical upper bound where all KVs fit in host memory.

Latency Distribution

End-to-End System Architecture

Bidaw integrates bidirectional awareness into both the compute engine (scheduler) and storage system (eviction policy). Metadata flows both ways to enable coordination.

Request Lifecycle

References & Related Work

Primary Paper: Bidaw: Enhancing Key-Value Caching for Interactive LLM Serving via Bidirectional Computation–Storage Awareness — Shipeng Hu et al., USENIX FAST'26
PDF
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng et al., 2023
Google Scholar
vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention — Woosuk Kwon et al., 2023
Google Scholar
Adaptive Replacement Cache (ARC) — Megiddo and Modha, 2003
Google Scholar
MiKV: No Token Left Behind — Mixed Precision KV Cache Compression, 2024-2025
Google Scholar
Related Episodes: Efficient KV Cache Reuse in Dynamic Agent Workflows
Listen
Related Episodes: 50x KV Cache Compression in Seconds via Attention Matching
Listen