AI Post Transformers • Interactive Visualization

IndexMem: Learned KV-Cache Eviction for Long-Context LLMs

This page treats the paper as a systems story in pictures: the KV cache grows with context, memory traffic starts to dominate, and IndexMem splits the fix into two distinct jobs, predicting what to keep and compensating for what gets evicted.

arXiv:2605.25475 Posted May 25, 2026 Xintong Yang et al. HKUST + Zhejiang University Transcript IDs: none beyond 2605.25475

Two problems, two modules

The selector estimates future token importance before those future queries fully exist. The latent memory handles the harder failure mode: once a token is evicted, can a compact state still return a useful correction later instead of pretending deletion was harmless?

Memory Anchor
7.68 → 4.08 GB
32K prefill + 1K decode efficiency run quoted in the episode.
Stress Test
56.0 vs 25.6
Llama 16K at 90% compression: IndexMem vs SnapKV.
Scope
3 backbones
Qwen3-8B, Mistral-7B-v0.3, Llama-3.1-8B-Instruct.
Caveat
Real problem / smaller proof
The evidence is bounded-cache long-context evaluation, not million-token deployment proof.

From prefill to the memory wall

Switch workload size to separate the part the paper measures from the part it mainly motivates.

Inference Pipeline

The cache starts as a decoding speedup and turns into a bandwidth tax as the working set grows.

Workload

Pressure vs Proof

As the prompt grows, the need becomes more urgent while the paper’s direct evidence covers a narrower slice.

KV Traffic Growth

Illustrative log-scale trend: the systems motivation grows faster than the measured region.

Full cache Evict + latent memory Extrapolation zone

Predicting future-relevant tokens

IndexMem’s selector is small on purpose: shared key projection, normalized scores, and max pooling so one future query can keep a quiet but crucial token alive.

Importance Heatmap

Hover cells to inspect future-query spikes. Compression acts on non-sink tokens while the first four sink positions stay pinned.

Query set
Eviction
Illustrative data, but the retention rule follows the episode’s logic: sink tokens remain, then max-over-queries ranks the rest.

Selector Mechanics

The important design choice is not “learned” by itself. It is the max pooling over queries, which preserves rare-but-critical evidence.

future query spike budget cutoff

Latent memory as compensation, not replay

Evicted keys and values write into a fixed-size state. Later queries read a summary vector back, and that vector only contributes a residual correction.

Write → Read → Residual

Use the step buttons to walk through the memory path. Toggle the correction off to see how much of the picture the selector alone must carry.

Mode
Step
surviving explicit tokens latent memory state query-dependent correction

What the paper shows under pressure

The episode’s reading is narrow by design: useful bounded-budget evidence, not a blanket proof of million-token readiness.

Benchmark Views

Switch between RULER stress points, LongBench degradation curves, and the one efficiency pass cited in the transcript.

View
Model
Bars and curves are anchored to transcript-quoted points where available, with realistic interpolation for the missing settings.

Needle Dead Zones

The episode calls out the visual pattern: fewer empty regions across needle positions when compression gets harsh.

References

Primary links only. arXiv links are rendered in canonical abs form when IDs were provided.