Dramatically improving long, multi-turn AI chatbot performance through bidirectional awareness between GPU scheduler and storage system.
Interactive LLM conversations require loading key-value tensors from all previous turns. With an average of 22 rounds per conversation, 93.1% of computation is redundant—recalculating previously computed KVs.
The GPU scheduler has no idea where KVs are stored (DRAM vs SSD), while the storage system doesn't know which conversations will heat up next. This coordination failure causes massive inefficiency.
GPU memory is too small. Host DRAM costs $5-10/GB. SSDs cost $0.10-0.20/GB but are 100-1000Ă— slower for random reads. Two-tier caching is essential, but coordination is the key.
Bidaw introduces two queues: a ready queue for requests with KVs in DRAM, and a preparing queue for requests loading from SSD. Ready requests never block on I/O.
Within the preparing queue, requests are reordered by estimated I/O time (primarily KV size). Smaller KVs load faster from SSD and get priority. Waiting time prevents starvation.
Traditional LRU fails because temporal locality is terrible—inter-arrival times are dominated by human think time. Bidaw predicts next access time by analyzing the model's response.
Bidaw combines answer length (longer = more reading time), question presence (prompts quicker reply), and sentiment (negative = slower re-engagement) into a weighted reuse distance. This is mapped to hit probability via a ghost cache.
Bidaw achieves up to 3.58Ă— latency reduction and 1.83Ă— throughput improvement, approaching the theoretical upper bound where all KVs fit in host memory.
Bidaw integrates bidirectional awareness into both the compute engine (scheduler) and storage system (eviction policy). Metadata flows both ways to enable coordination.