A visualization-first companion to the podcast episode: how RoPE rotates queries into position-specific frames, why recent post-RoPE attention can be a weak predictor for future importance, and how TriAttention uses pre-RoPE concentration plus trigonometric distance preferences to compress KV cache for long reasoning.
Systems win + mechanistic story. Mock overview chart reflects the paper’s headline positioning in the transcript.
Flow diagram: common recency-based policies observe a tiny recent post-RoPE window; TriAttention instead models head-specific distance preferences from pre-RoPE structure.
Hover cells. Same future-relevant token can look unimportant when judged only from a short recent post-RoPE window.
Different heads are not equally geometric. Concentrated heads benefit more from trig-based distance scoring; diffuse heads lean more on norms.
Toggle between pre- and post-RoPE. Queries at different positions occupy different orientations after rotation, so a short recent sample may be a poor predictor for later attention.
Mock MRL heatmap across heads × frequency bands. Hover for values. Bright bands indicate stronger angular concentration.
Step through the approximation: centers + RoPE imply a trigonometric attention-versus-distance prior.
Compare a recency-heavy policy to a TriAttention-like hybrid. Colored tokens indicate retained keys; hidden bars show whether the policy preserved future-useful positions.
Higher concentration ⇒ more weight on trigonometric distance term. Lower concentration ⇒ more fallback to norms.
Heads × token groups. Hover to inspect whether distant useful tokens survive under each policy.
Mock frontier inspired by transcript-reported results. Toggle the x-axis view to compare retained KV budget or relative memory reduction versus task accuracy.
Transcript numbers shown with approximate context: AIME24, AIME25, and MATH500 under compressed settings.
Memory cost scales with sequence length. Compression changes whether long reasoning fits on constrained hardware.