QVCache stores past query embeddings and their returned top-k results. A hit occurs only when a nearby cached query lies within that region’s learned acceptance radius.
Semantically similar queries rarely produce identical vectors. Toggle the mode to see why byte-identical keys miss the recurring intent.
Nearby-in-time queries can also be nearby in embedding space. The heatmap visualizes semantic closeness between query timestamps under different workload modes.
Each dot is a query in 2D embedding projection. Dashed windows indicate recent cacheable neighborhoods. Stronger locality produces overlapping top-k reuse opportunities.
A single global threshold over-accepts in unstable regions and under-accepts in stable ones. Click a region to compare local geometry and learned safe radius.
Thresholds shift as the system observes hit usefulness, misses, and backend overlap. The line chart shows how different regions converge to different reuse radii.
Mock data illustrates the paper’s core idea: a very fast hit path matters only if locality yields enough hits. Compare backend classes and cache hit rates.
Backend high-recall search pushes latency upward. QVCache aims to shift points leftward by bypassing repeated work while preserving backend-comparable recall on cache hits.
QVCache is not a new index. It lives above the retrieval engine and can complement graph, quantization, and disk-aware systems.
Cache hits bypass expensive retrieval before reranking or generation. Metadata filters and corpus updates still constrain safe reuse.
Should extra RAM go to backend expansion or semantic caching? The tradeoff depends on locality, backend slowness, and correctness constraints.