AI Post Transformers — Episode Companion

DeltaProduct: Extending DeltaNet's State-Tracking via Householder Products

Siems, Carstensen, Zela, Hutter, Pontil, Grazzi — NeurIPS 2025 arXiv:2502.10297 Freiburg · ELLIS Tübingen · Microsoft Research · Prior Labs · IIT · UCL

Why transformers and Mamba can't track state — and what changes

Transformers and diagonal linear RNNs (Mamba, GLA) sit inside a proven complexity ceiling — they cannot track arbitrary permutation-group state such as parity, no matter how much they scale. DeltaNet escapes this because its recurrence is secretly one step of gradient descent, which produces a Householder reflection as its state-transition matrix. DeltaProduct takes that idea further by chaining multiple reflections per token.

Left branch: architectures bounded by the TC0 circuit-complexity ceiling. Right branch: DeltaNet's gradient-descent reinterpretation, and DeltaProduct's generalization of it — explored in the next tab.

One reflection vs. two: why Cartan–Dieudonné matters

DeltaNet's update is I − β k kᵀ — a generalized Householder reflection derived from one gradient step on an associative-recall objective. DeltaProduct takes n_h steps per token, composing n_h reflections. Toggle below to see why going from one reflection to two is a qualitative jump, not an incremental one.

Dashed lines are the two mirror (reflection) planes. v₀ (cyan) is the initial state vector.

v₀ — start (0°) after reflection 1 (50°) after reflection 2 (rotation)
Cartan–Dieudonné theorem: any orthogonal transformation on n-dimensional space can be built from at most n reflections. Composing just two already yields a full rotation — orientation-preserving, unlike either reflection alone. No nonlinearity or readout sits between the two reflections inside one DeltaProduct step, so the composition is exact, not something depth has to learn to approximate.

Group word problems: S3, S4, A5, S5

Trained on sequences of 128 permutations, tested for extrapolation to 512. Color = generalization accuracy past training length at a single layer, for a given n_h. Hover any cell for the exact reading.

Illustrative data reflecting the paper's reported trends, not the exact reported numbers.

The surprise: naive reflection-counting predicts S4 needs n_h=3 and A5 needs n_h=4 — but both only need n_h=2. Why: S4 is isomorphic to the rotation group of a cube, A5 to the rotation group of a dodecahedron — both subgroups of SO(3), which only takes two reflections to generate a rotation. Learned betas cluster right at 2, and PCA on the keys shows 3 components explain over 95% of variance — the model found this structure with zero geometric priors.

Layers needed: DeltaNet (n_h=1) vs. DeltaProduct at optimal n_h

S5 never fits the training length under plain DeltaNet even at 10 stacked layers; DeltaProduct solves it in 1 layer at n_h=4 — the same reflection count the theorem predicts (n−1).

Real language models: 200M – 1.3B params, FineWeb

Length extrapolation is tested past the 4096-token training context on CodeParrot, OpenThoughts-Math, and TriviaQA. Switch the metric below to see why n_h=2 is the practical sweet spot.

n_h 1→2 produces a sharp improvement; n_h 2→3 nearly flattens. DeltaNet's state keeps accumulating rank past 4096 tokens, drifting out of distribution; gated DeltaProduct heads learn to reset and bound it. Illustrative data reflecting the paper's reported trends.

Does the gap survive scale?

Parameter-matched by scaling head dimension (Figure 10 in the paper). DeltaProduct (n_h=2) keeps its edge to the largest scale tested, 1.3B params — the recurrence itself costs n_h× the sequential compute per token, clawed back partly by a Triton kernel ~20% faster than DeltaNet's baseline.

The n_h dial: expressivity vs. compute

2
2× sequential compute vs DeltaNet
✗ S3 ✗ S4 ✗ A5 ✗ S5

Move the slider to see which group word problems become solvable in a single layer at each n_h, against the linear compute cost it buys.

Architecture comparison

MetricDeltaNet (n_h=1)DeltaProduct n_h=2DeltaProduct n_h=3RWKV-7
State-tracking ceiling Diagonal + rank-1; S3/S4/A5 fail past training length S3, S4, A5 solved; S5 fails S3, S4, A5 solved; S5 improves, not solved Proven: any regular language, 4 layers (theory)
Sequential compute / token 1×2×3×Variable — no head-to-head run in this paper
Spectral norm bound ≤ 1 (stable)≤ 1 (stable)≤ 1 (stable) Not bounded — trades stability for per-layer expressivity
Largest scale tested 1.3B params, FineWeb only Not directly compared here

What this paper doesn't show

Synthetic-only state-tracking evidence. Every S3/S4/A5/S5 result comes from one benchmark family (Merrill, Petty & Sabharwal, ICML 2024), trained at length 128, tested to 512 — not arithmetic traces, code interpreters, or agent tool-call chains.
Sub-1.3B scale only. All language-model results stop at 1.3B params on FineWeb; architectural gaps between RNNs and Transformers have flipped before once real scale arrived.
No compute-matched comparison. Figure 10 is parameter-matched, not FLOPs-matched — since DeltaProduct's recurrence cost scales linearly with n_h, a matched parameter count is quietly a bigger training budget.
No RWKV-7 head-to-head. RWKV-7 sits in the paper's own comparison table as the closest theoretical rival, but there's no empirical run against it on language modeling or state-tracking.
Correlation, not causation, on effective rank. The paper's own phrase is "we attribute" — a hedge. A causal test would clamp DeltaNet's rank artificially and see if the loss curves move.
Theorem 1's hardest construction is never built. The lookup-table argument showing 4 layers at n_h=1 can in principle solve any group word problem is never actually implemented or trained.

Verdict: an elegant mechanism, tied to a real theorem, that measurably buys expressivity DeltaNet lacks — with the S4/A5 isomorphism a genuine surprise. Good NeurIPS-caliber science on a specific mechanism; not a settled verdict on linear RNNs beating Transformers at scale.

References

  1. 1. DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products — J. Siems, T. Carstensen, A. Zela, F. Hutter, M. Pontil, R. Grazzi, 2025
  2. 2. Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues — R. Grazzi, J. Siems, A. Zela, J. Franke, F. Hutter, M. Pontil, 2025 (ICLR)
  3. 3. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — S. Yang, B. Wang, Y. Zhang, Y. Shen, Y. Kim, 2024 (NeurIPS)
  4. 4. Gated Delta Networks: Improving Mamba2 with the Delta Rule — S. Yang, J. Kautz, A. Hatamizadeh, 2025 (ICLR)
  5. 5. RWKV-7 "Goose" with Expressive Dynamic State Evolution — B. Peng, R. Zhang, D. Goldstein, et al., 2025
  6. 6. The Illusion of State in State-Space Models — W. Merrill, J. Petty, A. Sabharwal, 2024 (ICML)
  7. 7. Fixed-Point RNNs: From Diagonal to Dense in a Few Iterations — S. Movahedi, F. Sarnthein, N. Muca Cirone, A. Orvieto, 2025