AI Post Transformers · visual companion

QVCache for Semantic Caching in ANN Search

A query-aware semantic cache in front of ANN backends: visualize temporal-semantic locality, region-specific reuse thresholds, and the latency/recall tradeoff of bypassing repeated high-recall vector search.
backend-agnostic cache layer
sub-ms hit path
bounded memory
recall@k-sensitive reuse
paper
Göçer et al. · 2026
core pattern
temporal-semantic locality
acts on
query embeddings + top-k cache
fits with
HNSW · PQ · DiskANN · FAISS

Request Path: cache in front of any ANN backend

QVCache stores past query embeddings and their returned top-k results. A hit occurs only when a nearby cached query lies within that region’s learned acceptance radius.

query embedding cache mini-index ANN backend fast hit path

Step-through: exact-match cache vs semantic cache

Semantically similar queries rarely produce identical vectors. Toggle the mode to see why byte-identical keys miss the recurring intent.

Hover query circles and cache cells for distances, overlap, and hit decisions.

Temporal-Semantic Locality Matrix

Nearby-in-time queries can also be nearby in embedding space. The heatmap visualizes semantic closeness between query timestamps under different workload modes.

Cluster stream and reuse windows

Each dot is a query in 2D embedding projection. Dashed windows indicate recent cacheable neighborhoods. Stronger locality produces overlapping top-k reuse opportunities.

cluster A cluster B cluster C active cache windows

Region-specific thresholds

A single global threshold over-accepts in unstable regions and under-accepts in stable ones. Click a region to compare local geometry and learned safe radius.

Online adaptation loop

Thresholds shift as the system observes hit usefulness, misses, and backend overlap. The line chart shows how different regions converge to different reuse radii.

Latency vs recall across backends

Mock data illustrates the paper’s core idea: a very fast hit path matters only if locality yields enough hits. Compare backend classes and cache hit rates.

Recall / speed frontier

Backend high-recall search pushes latency upward. QVCache aims to shift points leftward by bypassing repeated work while preserving backend-comparable recall on cache hits.

Index families

QVCache is not a new index. It lives above the retrieval engine and can complement graph, quantization, and disk-aware systems.

Where it fits in RAG serving

Cache hits bypass expensive retrieval before reranking or generation. Metadata filters and corpus updates still constrain safe reuse.

Memory budget question

Should extra RAM go to backend expansion or semantic caching? The tradeoff depends on locality, backend slowness, and correctness constraints.

References

  1. Göçer, Tsakalidou, Nicholson, Kim, Ailamaki. QVCache: A Query-Aware Vector Cache.
  2. Arefin et al. Survey on Nearest Neighbor Search Methods.
  3. Jégou, Douze, Schmid. Product Quantization.
  4. Malkov, Yashunin. HNSW.
  5. Subramanya et al. DiskANN.
  6. Johnson, Douze, Jégou. FAISS.
  7. Bratseth et al. Vespa.
  8. Andrew Kane. pgvector.
  9. Milvus team. Milvus.
  10. Freund, Schapire. On-Line Learning.
  11. Duchi, Hazan, Singer. AdaGrad.
  12. Luo et al. Semantic Caching.
  13. GPTCache contributors. GPTCache.
  14. Chen et al. SPANN.