A visual companion focused on the scheduling trick: keep dense attention on GPU, let CPU scout only a sparse offloaded slice, and overlap pre-compute, transfer, and decode so long-context serving stops idling on memory traffic.
Switch between a stall-heavy recall baseline and ScoutAttention’s overlapped schedule.
Baseline offloading treats CPU memory as a bigger closet, then waits for data to come back. ScoutAttention changes the schedule instead: GPU stays on dense resident blocks while the CPU prepares a small offloaded candidate set one layer ahead.
The paper’s reported asymmetry matters. If GPU attention is about 20× faster than CPU attention, the CPU cannot be an equal co-worker. It has to be narrow, early, and hidden under work that would happen anyway.
Hover the rows in the SVG to inspect stage timing and where stalls appear.
Use the workload toggle and step buttons to watch which blocks matter and how stable the shortlist remains.
These values are synthetic but shaped to match the episode’s claims: ScoutAttention improves throughput while keeping quality loss modest.
The strongest win appears when context is long and locality stays stable for several nearby decode steps.
Retrieval jumps, code-symbol land mines, or sudden mode switches can reshuffle important blocks fast. Then the CPU’s prepared shortlist is partly wrong, periodic recall fires more often, and tail latency can widen even if the mean still looks good.
That is why comparisons against eviction-focused methods and per-regime reporting matter. “Average degradation” can hide exactly the cases where one missed block ruins the answer.
Hover cells in the heatmap to inspect predicted win and risk level.
Compact links for the main paper, adjacent systems work, and prior podcast episodes.