Mock comparative scores summarize the conversation’s framing: old linear attention often lost on quality and hardware efficiency; GLA tries to improve both.
Failure modes called out: weak model quality and weak real-kernel efficiency.
Paper claims standalone layer speed can beat FlashAttention-2 already at modest length.
Train-short / test-long narrative highlights usable long-context generalization.
Click models to inspect their memory style. The diagram is intentionally spatial: left/right tracks explicit retrieval versus recurrent compression; up/down tracks static decay versus selective data-dependent control.
Use the stepper to watch a toy sequence fill memory. The heatmaps show a compressed matrix-state evolving over time; gating can keep, decay, or overwrite features instead of blindly accumulating them.
Toggle between a naive implementation and flash-style chunked training. The Sankey-like flow and heatmap show how reducing writes to HBM can dominate whether a “linear-time” idea becomes genuinely fast on GPUs.
These mock-but-realistic charts encode the episode’s argument: beating old linear attention is not enough; the comparison target is optimized softmax attention such as FlashAttention-2, plus adjacent recurrent contenders.