Fixed-Depth vs. Looped Architecture
A normal transformer stacks N distinct blocks, each with its own weights — a token passes through each exactly once. A looped model builds one block and replays it through the residual stream L times, training with a shallow L and testing with a deeper one, at constant parameter count.
Why this matters: RNN over depth, not sequence
The useful analogy isn't "bigger transformer" — it's an RNN. An RNN applies the same cell repeatedly across the sequence dimension. A looped transformer applies the same block repeatedly across the depth dimension. Same weights, revisited — which is exactly why the dynamical-systems lens (recurrence math) applies.
The Transition Matrix A: Does the Stream Blow Up?
Write the loop update as h(t+1) = A·h(t) + B·x + nonlinearity. Drop the nonlinear term and you get a clean LTI system — stability is governed entirely by the spectral norm ρ(A). Above 1, repeated looping explodes exponentially; below 1, it stays bounded.
Additive injection forces A to the identity matrix — ρ(A)=1 exactly, riding the edge without decaying or exploding on its own.
ρ(A) Across Training Steps
Table 1 / Figure 3 from the paper, reconstructed: runs where ρ(A) crosses above the instability threshold (dashed line) are exactly the runs whose loss spikes shortly after. Parcae's ZOH parameterization keeps ρ(A) provably under the line by construction.
isoFLOP Curves: Recurrence as a Third Scaling Axis
Chinchilla-style: hold total FLOPs fixed, sweep mean recurrence depth against token count, fit a parabola per budget. The minima trace a power law — recurrence is a real, plannable scaling axis, not an arbitrary knob.
Fitted Power-Law Exponents (log‑log)
Optimal recurrence scales roughly as FLOPs0.40; optimal token count as roughly FLOPs0.78 — consistent across the 140M and 370M model sizes tested. Note: unverified beyond these two points.
Quality: Parcae vs. Transformer vs. RDM
Parcae 770M matches the Core score of a 1.3B parameter Transformer — roughly 87.5% relative quality at half the parameters. All comparisons here use the shared nanochat / FineWeb-Edu setup (not the separate Huginn perplexity track).
Test-Time Loop Scaling: A Saturating Ceiling
More loop steps at inference help fast, then flatten — the floor is set by how much recurrence the model saw during training. "Train small, dial up compute at inference" mostly evaporates once you exceed that training-time depth.