Interactive Visualization • Systems for Long-Context LLM Inference

ScoutAttention for Efficient KV Cache Offloading

A visual companion focused on the scheduling trick: keep dense attention on GPU, let CPU scout only a sparse offloaded slice, and overlap pre-compute, transfer, and decode so long-context serving stops idling on memory traffic.

Critical path: stop making recall the event that gates every layer

Switch between a stall-heavy recall baseline and ScoutAttention’s overlapped schedule.

GPU dense attention / execution CPU sparse scoring / shortlist PCIe or recall traffic Visible stall on decode critical path

What changes

Baseline offloading treats CPU memory as a bigger closet, then waits for data to come back. ScoutAttention changes the schedule instead: GPU stays on dense resident blocks while the CPU prepares a small offloaded candidate set one layer ahead.

The paper’s reported asymmetry matters. If GPU attention is about 20× faster than CPU attention, the CPU cannot be an equal co-worker. It has to be narrow, early, and hidden under work that would happen anyway.

Hover the rows in the SVG to inspect stage timing and where stalls appear.

Block forecast stability is the bet

Use the workload toggle and step buttons to watch which blocks matter and how stable the shortlist remains.

Rows are layers. Columns are KV blocks from newest on the left to cold history on the right. Dashed outline marks GPU-resident blocks; halo marks the CPU shortlist prepared one layer ahead.
The point is not future peeking. The CPU uses current hidden-state signals and recent attention locality to pre-score offloaded blocks for the next layer.

Mock result board for three serving regimes

These values are synthetic but shaped to match the episode’s claims: ScoutAttention improves throughput while keeping quality loss modest.

The comparison frame: full attention, recall-heavy offload, CPU co-attention, and ScoutAttention across chat, RAG, and long-reasoning workloads.
Upper-right is the desirable region: higher throughput, higher quality retention. ScoutAttention aims to move there by turning blocking transfer into overlapped background work.

Where it helps, and where it can go stale

The strongest win appears when context is long and locality stays stable for several nearby decode steps.

Benefit heatmap across context length and prompt volatility. Hot cells mean better overlap payoff; cooler cells mean the shortlist or recall cadence can go stale faster.

Pressure points

Retrieval jumps, code-symbol land mines, or sudden mode switches can reshuffle important blocks fast. Then the CPU’s prepared shortlist is partly wrong, periodic recall fires more often, and tail latency can widen even if the mean still looks good.

That is why comparisons against eviction-focused methods and per-regime reporting matter. “Average degradation” can hide exactly the cases where one missed block ruins the answer.

Hover cells in the heatmap to inspect predicted win and risk level.

References

Compact links for the main paper, adjacent systems work, and prior podcast episodes.