Standard attention keeps an expanding token-level memory. Linear attention and state-space variants compress the past into a fixed state. This is GPU-friendly, but retrieval quality depends on how that state is updated.
Mock curves illustrating why the field tolerates compressed memory despite imperfect recall.
Each token writes a key-value association into a shared finite state. Hover cells to inspect interference hotspots.
Retrieval-heavy tasks stress whether the memory rule stores and replaces associations cleanly. The heatmaps below simulate how broad gating, delta overwrite, and the combined rule change retrieval sharpness as more items are packed into a fixed-size state.
Gating is a coarse retain/erase knob. Delta updates are surgical corrections tied to a key. The combined layer uses both: broad forgetting for context shifts, targeted replacement for memory editing.
Mock benchmark summaries based on the episode’s qualitative pattern: pure Gated DeltaNet improves over parent baselines, while hybrids often push further. The strongest scientific claim is “strong ingredient in a strong recipe,” not necessarily sole cause.
Hybrids mix exact local attention with compressed global memory. They often score best, but attribution becomes messier.
How much of the observed gain plausibly comes from the update rule, local attention, implementation quality, or larger memory budget?