MoBA routes each query to a few key-value blocks instead of the full sequence. This page turns the episode into a visual lab: block routing geometry, the sqrt(d / B) signal-to-noise law, FlashMoBA’s gather-densify-scatter kernel, and the tradeoff surface between retrieval quality and hardware efficiency.
Queries score block centroids, select top-k blocks, then run dense attention only where the router points. The whole page starts here because every later chart depends on this compression step.
A block centroid is a mean over B keys. If only one token matters, the query is betting that one high-affinity key survives averaging against B-1 distractors.
Hover the matrix. Higher values indicate cleaner separation between relevant and irrelevant blocks. Toggle convolution to increase effective signal clustering inside each block.
Switch between kernel-centric speed and training-quality views. The episode’s headline is not just “small blocks help”; it is “small blocks help only if the kernel keeps the GPU busy.”
The sparse routing pattern is reorganized into batches of dense work per selected block. This avoids materializing the full query-by-centroid score matrix and reduces global-memory churn.
The episode pushes hard on what the paper does not show: frontier-scale validation, router stability, positional interactions, and downstream capability retention.
A viable production path probably blends small-block routing with MoE-style guardrails, positional ablations, and hybrid sparse+dense designs rather than treating MoBA as all-or-nothing.