LinearKV: When Exact State Merging Breaks Hybrid Models

One cached state suffices for position-independent caching in hybrid Mamba-attention LLMs — and the mathematically "exact" way to merge cached recurrent state degrades one tested model's quality by more than half, while a simpler shortcut wins.

arXiv:2608.11231 Interactive Viz ↗ Liu, Qi, Wang et al. · TeleAI / SJTU / XJTU / U. Buffalo · 2026

Decoupled Initialization — where full attention and recurrence split

Full-attention layers still concatenate cached KV in order. Recurrent layers (Mamba-2, Gated DeltaNet) need a function f to fold K cached local states into one initial state before recompute. Toggle which f is highlighted below.

Exact composition algebraically chains all K cached transitions to reconstruct the "true" full-prefix state — but each chunk's cached transition was built from a chunk prefilled in isolation, so composing them is exact only relative to a flawed reference.

Why Recurrent Layers Can't Just "Patch a Few Tokens"

Full-attention position-independent caching repairs a cache by recomputing a small slice of tokens against a concatenated KV ledger. Recurrent layers keep no token-indexed ledger — just one fixed-size state — so there's nothing granular left to patch.

State Error by Layer Depth (Figure 3 reproduction, mock values)

Relative state-error norm vs. layer index. GDN models (OLMo, Qwen) stay under 1.0 at every layer; the Mamba-2 model (Granite) compounds error, spiking to ~2x the state norm at one deep layer. Hover a point for its value.

2.0x
Granite peak error / state norm
<1.0
OLMo & Qwen (GDN), all layers
0.01
Layer-0 floor (bf16 rounding), all models

Why the Transition Rule Decides the Outcome

Mamba-2's transition is a bare scalar decay — it can only retain and sum error forward. GDN's is a dense, gated, direction-dependent correction — it can suppress specific error directions. Mock 6×6 transition-effect matrices below (hover cells).

Mamba-2: scalar decay (uniform)

Gated DeltaNet: dense gated correction

Error recursion (Eq. 7): error at chunk j = transition(previous error) + freshly injected mismatch. A uniform scalar transition only accumulates; a dense gated transition can damp specific directions to near zero.

Quality Recovered at 20% Recompute Budget (EPIC selector)

% of full-context quality recovered. On the Mamba-2 model (Granite), exact composition collapses; on the two GDN models it's a wash.

Budget Sweep on Granite — the gap never closes

Sweeping recompute budget 3%–40% while holding selector and positions fixed. Toggle metric to see quality vs. time-to-first-token.

Exact composition stays flat at 41–52% of full quality across every budget from 3% to 40%. Single-block reaches 76–89%. The initializer sets a quality ceiling recompute can't buy back.

Prefix Caching vs. Position-Independent Caching

Toggle between the two caching regimes to see why chunk order matters for one and not the other.

Lineage: Marconi → HYPIC & LinearKV

Marconi solved prefix reuse for hybrids. HYPIC and LinearKV are concurrent 2026 work racing to solve position-independent reuse for hybrids — and they disagree on the initializer.

References