AI Post Transformers · Episode Companion

Parcae: Stabilizing Looped Language Models with Control Theory

Instead of stacking distinct transformer layers, Parcae applies one block repeatedly through the residual stream. This page visualizes the paper's core move: reframing looped inference as a linear time-invariant dynamical system, where the spectral norm of transition matrix A predicts — and a Mamba-style parameterization guarantees — whether the stream stays bounded or explodes.

arXiv:2604.12946 Prairie, Novack, Berg-Kirkpatrick, Fu · 2026 UC San Diego · Together AI

Fixed-Depth vs. Looped Architecture

A normal transformer stacks N distinct blocks, each with its own weights — a token passes through each exactly once. A looped model builds one block and replays it through the residual stream L times, training with a shallow L and testing with a deeper one, at constant parameter count.

shared / repeated weights distinct per-layer weights injected prelude embedding

Why this matters: RNN over depth, not sequence

The useful analogy isn't "bigger transformer" — it's an RNN. An RNN applies the same cell repeatedly across the sequence dimension. A looped transformer applies the same block repeatedly across the depth dimension. Same weights, revisited — which is exactly why the dynamical-systems lens (recurrence math) applies.

The Transition Matrix A: Does the Stream Blow Up?

Write the loop update as h(t+1) = A·h(t) + B·x + nonlinearity. Drop the nonlinear term and you get a clean LTI system — stability is governed entirely by the spectral norm ρ(A). Above 1, repeated looping explodes exponentially; below 1, it stays bounded.

Spectral norm ρ(A)
1.00
Stability verdict
Marginal

Additive injection forces A to the identity matrix — ρ(A)=1 exactly, riding the edge without decaying or exploding on its own.

ρ(A) Across Training Steps

Table 1 / Figure 3 from the paper, reconstructed: runs where ρ(A) crosses above the instability threshold (dashed line) are exactly the runs whose loss spikes shortly after. Parcae's ZOH parameterization keeps ρ(A) provably under the line by construction.

isoFLOP Curves: Recurrence as a Third Scaling Axis

Chinchilla-style: hold total FLOPs fixed, sweep mean recurrence depth against token count, fit a parabola per budget. The minima trace a power law — recurrence is a real, plannable scaling axis, not an arbitrary knob.

Fitted Power-Law Exponents (log‑log)

Optimal recurrence scales roughly as FLOPs0.40; optimal token count as roughly FLOPs0.78 — consistent across the 140M and 370M model sizes tested. Note: unverified beyond these two points.

optimal recurrence ∼ FLOPs^0.40 optimal tokens ∼ FLOPs^0.78

Quality: Parcae vs. Transformer vs. RDM

Parcae 770M matches the Core score of a 1.3B parameter Transformer — roughly 87.5% relative quality at half the parameters. All comparisons here use the shared nanochat / FineWeb-Edu setup (not the separate Huginn perplexity track).

Transformer RDM Parcae

Test-Time Loop Scaling: A Saturating Ceiling

More loop steps at inference help fast, then flatten — the floor is set by how much recurrence the model saw during training. "Train small, dial up compute at inference" mostly evaporates once you exceed that training-time depth.

References