AI Post Transformers May 13, 2026 Interactive Visualization

Long Context Pre-Training with Lighthouse Attention

A systems-first bet on cheap long-context pretraining: route tokens through a hierarchical sparse selector, keep the expensive inner compute inside ordinary dense FlashAttention, then switch back to dense attention late in training and ask whether the final model still behaves like a real dense transformer.

Paper Metadata

Dense kernel retained Hierarchical routing Preliminary evidence

Episode Snapshot

Context
98,304
Model
530M
Layers
26 sparse / 4 dense
Claim
Sparse-first → dense recovery

Dense Attention Wall vs Lighthouse Route

Same transformer family, different routing geometry. The left side shows why long-context dense attention blows up; the right side shows the paper’s wrapper: hierarchical scoring outside, ordinary FlashAttention inside.

Quadratic Pressure
Dense token-to-token map expands as N²
Key Bet
Move selection outside the dense kernel
Kernel Strategy
Reuse FlashAttention on gathered subsequence
Final Target
Ordinary dense serving model after resumption

Step-by-Step Relay

The method is easiest to see as a five-stage relay race. Each step reveals a different slice of the sparse wrapper around a dense core.

Multi-Scale Selector Heatmaps

Mock attention-routing grids illustrate the paper’s hierarchical search idea: coarse pooled views prune the search space before a dense kernel runs on a smaller causal subsequence.

low relevance candidate region top-K selected
Hover cells to inspect pooled blocks, selected regions, and per-stage sparsity.

Scaling Economics

FlashAttention reduces waste, but not the quadratic law. These mock curves show why a sparse-first pretraining schedule becomes tempting once context length moves from 8K toward 128K and beyond.

Dense SDPA
Exact, hardware-optimized, still quadratic
Lighthouse
Approximate routing plus dense inner kernel
Takeaway
Systems gain depends on gather/scatter not dominating

Recovery Schedule and Evidence Quality

The strongest claim is not “sparse attention is good.” It is “train mostly sparse, switch back to dense, and recover dense-model quality.” The chart below separates what the paper suggests from what remains unproven.

References and Extracted arXiv IDs

Transcript arXiv-pattern hits: 2605.06554
Long Context Pre-Training with Lighthouse Attention
arXiv:2605.06554
FlashAttention
Scholar search
H-Transformer-1D
Scholar search
Native Sparse Attention
Scholar search
Every Token Counts
Scholar search