1 00:00:01,000 --> 00:00:39,024 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called Parcae: Scaling Laws For Stable Looped Language Models, from Hayden Prairie et al. — that's three co-authors alongside him, Zachary Novack, Taylor Berg-Kirkpatrick, and Daniel Y. Fu — out of University of California, San Diego, with a Together AI affiliation as well. It went up on arXiv on April 14th, 2026. And Ada, I'll be honest, when I saw the phrase 'dynamical system' in the abstract my brain immediately filed this under 'skip to the results.' 2 00:00:39,024 --> 00:01:10,000 [Dr. Ada Shannon] Don't skip, Hal, this is the good stuff. Most looped-model papers I read are basically vibes — someone loops a transformer block a few extra times, quality goes up a little, they ship a leaderboard number and call it a day. What actually pulled me in here is that they don't just report it works, they explain why the previous attempts kept breaking, using honest-to-god control theory. That's rare. Usually when a paper says 'we recast the problem as a dynamical system,' it's decoration. Here it's load-bearing — the stability fix falls directly out of the math, not out of a hyperparameter sweep somebody got lucky with. 3 00:01:10,000 --> 00:01:49,799 [Hal Turing] Okay, so let's back up for anyone who hasn't touched this corner of the field, because I hadn't either until this week. A normal transformer — GPT, Llama, whatever — is what you'd call fixed-depth. You stack N distinct blocks of layers, each with its own separate weights, and a token flows through each block exactly once, start to finish. Want a better model? You usually add more blocks, meaning more parameters, or you throw more training data at it. A looped model throws that out. Instead of N different blocks, you build one block and you run the residual stream through it over and over — say, four times during training, and maybe twelve times at inference if you want the model to 'think' harder. 4 00:01:49,799 --> 00:02:32,050 [Dr. Ada Shannon] Right, and the useful comparison is to an RNN, not to a bigger transformer. A recurrent net applies the same cell repeatedly across the sequence dimension — token by token. A looped transformer applies the same block repeatedly across the depth dimension — the same weights, revisited. Same parameter count no matter how many loops you run, which is the entire appeal. If you're deploying on a phone or some edge box with a tight memory budget, you'd love GPT-4-class reasoning depth out of something with the memory footprint of a much smaller model. That's the pitch. This isn't new, either — Dehghani and colleagues proposed Universal Transformers back in 2018, and ALBERT tied weights across BERT's layers in 2019 to shrink its footprint. The idea's been sitting there for years. 5 00:02:32,050 --> 00:02:37,950 [Hal Turing] So if it's been around that long, why isn't every model looped right now? What's the catch? 6 00:02:37,950 --> 00:03:21,000 [Dr. Ada Shannon] Instability. Reapplying the same nonlinear block dozens of times is mathematically just iterating a dynamical system, and if that system's transition amplifies the signal even slightly, the residual stream doesn't just drift — it can blow up exponentially with loop count. In practice that shows up as loss spikes, sometimes total divergence mid-training. The paper points directly at RDM, the Recurrent Depth Model line from Geiping and colleagues' 2025 work on latent reasoning — that's the prior looped baseline everyone's been building on, and it needed careful residual normalization and touchy hyperparameter tuning just to stay upright. Most labs looked at that fragility and went back to fixed-depth scaling, where Kaplan-style and Chinchilla-style laws are boringly predictable. 7 00:03:21,000 --> 00:03:25,750 [Hal Turing] Which brings us to the dynamical systems angle you were teasing— 8 00:03:25,750 --> 00:04:05,650 [Dr. Ada Shannon] —exactly, hold that thought, because this is the actual contribution. Take the loop update and write it the way control theorists write any evolving state: h at the next step equals A times h now, plus B times the input, plus whatever nonlinear mess the transformer block adds on top. Drop that nonlinear mess for a second and you're left with a clean linear time-invariant system — an LTI system. And for an LTI system there's a completely standard, decades-old answer to 'does this blow up.' It depends entirely on the spectral norm of A. Above one, it explodes under repeated iteration. Below one, it's bounded. That single number turns a mysterious training failure into something you can actually check. 9 00:04:05,650 --> 00:04:13,125 [Hal Turing] So instead of 'the loss spiked, who knows why,' you get an actual measurable culprit sitting inside the weights. 10 00:04:13,125 --> 00:04:45,524 [Dr. Ada Shannon] That's the reframe. And once you've got a model that doesn't explode, you can ask a much more interesting question: does looping behave predictably as you scale it? That's where the paper leans on isoFLOP fitting — the same Chinchilla-style trick of holding total training compute fixed and sweeping the allocation, except instead of just splitting between parameters and data, they're adding a third knob: how many times you loop. Fit parabolas across a bunch of fixed-FLOP budgets, find the minimum, and you get an empirical curve for the FLOP-optimal recurrence depth. We'll get into what that curve actually looks like next. 11 00:04:45,524 --> 00:05:11,024 [Hal Turing] Hold on, before we get to the actual scaling curves, how did they prove that? "The residual stream blows up" is the symptom, but what's the actual measurable cause? Because an RDM training run just looks like any other exploding-loss run from the outside, the loss chart goes vertical and you shrug and lower the learning rate. What did the Parcae authors point to that says "here, this specific number sitting in the weights, is why it's happening"? 12 00:05:11,024 --> 00:05:53,500 [Dr. Ada Shannon] They point at the eigenvalues of the A matrix in that dynamical-systems formulation we just built. Table 1 in the paper lays it out cold: inject the prelude embedding by addition, the way some prior looped models do, and A comes out as the identity matrix, so its spectral norm rho of A equals exactly 1, marginally stable, riding the edge without decaying or exploding on its own. Inject by concatenation with a learned projection, which is what Geiping's RDM does, and A is an unconstrained learned matrix, so rho of A can land anywhere, including well above 1, unstable by definition. Figure 3 shows this isn't theoretical: they track rho of A across training steps at different learning rates, and every run that diverges is one where rho of A crosses above 1 right before the loss spikes. 13 00:05:53,500 --> 00:06:21,000 [Hal Turing] Oh wait wait wait, that's the same move as Mamba, isn't it? You force the diagonal state matrix into a provably stable region through the parameterization itself, instead of hoping gradient descent stumbles into it. That's exactly how Gu and Dao keep S4 and Mamba's state-space recurrence from blowing up over long sequences, negative real parts on the diagonal, guaranteed decay. Are they literally borrowing that trick here, just pointed at loop depth instead of sequence position? 14 00:06:21,000 --> 00:07:05,324 [Dr. Ada Shannon] Exactly right, and it's a sharp piece of borrowing. Parcae parameterizes A as a negative diagonal matrix, then discretizes it with a zero-order hold, the same ZOH move from Dao and Gu's "Transformers are SSMs" paper, Stanford and CMU, 2024. The negative-exponential form guarantees every diagonal entry is negative, so after discretization rho of A is provably under 1 by construction, not by luck. Same idea as Mamba, just applied along the loop-depth axis instead of the sequence-time axis. Then two more stabilizers get layered on: a normalization layer on the injected embedding e, which fixes late-training loss spikes they saw appear around 170,000 steps, and per-sequence depth sampling, drawing a different recurrence depth for each sequence within a micro-batch instead of one depth for the whole batch. 15 00:07:05,324 --> 00:07:29,300 [Hal Turing] Okay, that's a clean fix on paper, but does it hold up in practice, or is this one of those stability patches that stops the loss from exploding and then quietly tanks quality because you've handcuffed the architecture? Constraining eigenvalues sounds great until you realize you might've just made the model less expressive to buy that stability. So what happens when they actually run it at scale? 16 00:07:29,300 --> 00:08:17,925 [Dr. Ada Shannon] It holds up, and that's the interesting part. Against parameter- and data-matched RDMs, Parcae cuts validation perplexity by up to 6.3 percent. Against parameter-matched Transformers, same nanochat setup, FineWeb-Edu data, Parcae gains up to 2.99 points on Core and 1.18 on Core-Extended. The number that actually stopped me: their 770-million-parameter Parcae matches the Core score of a 1.3-billion-parameter Transformer, roughly 87.5 percent relative quality of a model twice its size. One methods detail worth flagging now, we'll come back to it: they didn't run a fresh hyperparameter sweep for Parcae at all. They reused the sweeps already tuned for the RDM and Transformer baselines and dropped Parcae's runs straight into those same settings. 17 00:08:17,925 --> 00:08:42,600 [Hal Turing] So it's winning without even getting its own tuning budget, which reads as either impressive or a red flag depending on how you squint at it, we'll get into that properly later. But you teased isoFLOP scaling laws a minute ago before I derailed you with the Mamba tangent. Now that the model doesn't fall over, what does that sweep actually show once you're free to trade recurrence against data at a fixed FLOP budget? 18 00:08:42,600 --> 00:09:36,100 [Dr. Ada Shannon] They train 140-million and 370-million Parcae models at fixed FLOP and parameter budgets, sweeping token count against mean recurrence, then fit a parabola to each FLOP budget the way Hoffmann fit Chinchilla. The minima trace out power laws: optimal recurrence scales roughly as FLOPs to the 0.40, optimal token count as roughly FLOPs to the 0.78, consistent across both model sizes. So recurrence is a real third scaling axis alongside parameters and data, not just a knob you crank arbitrarily. Separately, at test time, they sweep how much looping helps at inference and find a saturating exponential decay, quality improves fast with more loop steps, then flattens near a floor set by how much recurrence the model saw during training. They fold both into one unified equation, where the training law sets the floor and a decay term captures how fast extra test-time loops close the gap to it. 19 00:09:36,100 --> 00:10:08,000 [Hal Turing] Okay, gamma-mu around 0.40 for recurrence, gamma-D around 0.78 for tokens — clean power laws. But the fit and the held-out validation both happen at the same two sizes, 140 million and 370 million. Nobody's tested whether it holds at 770 million or 1.3 billion. That's exactly the position Chinchilla was in before people reran Hoffmann's fits and found the exponents needed revising. Genuine scaling law, or a confident two-point line? 20 00:10:08,000 --> 00:10:29,925 [Dr. Ada Shannon] A bit of both, credit where due — they admit it themselves, buried in the limitations section: it 'remains to be seen' if Parcae holds at large FLOP budgets. Unusually honest for a scaling-laws paper. But it sits oddly next to a title that claims 'scaling laws' unqualified, and an abstract leaning on edge deployment as motivation, when everything here tops out at 1.3 billion parameters and roughly a hundred billion tokens — nowhere near production scale. 21 00:10:29,925 --> 00:10:53,325 [Hal Turing] Here's the tuning-budget point we flagged earlier, and it cuts both ways. Does zero-sweep Parcae beating two fully-tuned baselines make the win stronger, since it's outperforming tuned opponents cold? Or does it quietly undercut Table 2's claim that Parcae is 'more robust to hyperparameter selection' when it was never actually tested outside the range those baselines needed? 22 00:10:53,325 --> 00:11:12,500 [Dr. Ada Shannon] Both readings survive at once, honestly. It's genuinely impressive that a zero-tuning model beats two separately-tuned baselines — that's the strongest version of the robustness story. But 'more robust' is a claim about behavior outside the tested range, and they never left the range the baselines happened to nee— 23 00:11:12,500 --> 00:11:38,850 [Hal Turing] —Sorry, cutting you off, but that same instinct applies to the headline number too. The 6.3 percent perplexity drop over RDMs is measured at 100 million and 350 million, on the Huginn dataset and tokenizer. Every Core benchmark number — the 140 million to 1.3 billion results — runs on a completely different setup, nanochat and FineWeb-Edu. Two non-overlapping tracks, cited in the same breath. 24 00:11:38,850 --> 00:12:20,100 [Dr. Ada Shannon] Exactly, and dataset, tokenizer, and training recipe all move perplexity independently of architecture, so there's no clean way to isolate how much of that 6.3 percent is the stability fix versus just Huginn being a different corpus. I'd trust the Core numbers more, since Transformer and Parcae share a setup there. There's a second ceiling too — Figure 7 confirms that ceiling sits right at the training-time mu-rec, so 'train small, dial up compute at inference' mostly evaporates once you can't exceed what you trained for. And every quality claim rests on Core, Lambada, Hellaswag, ARC, PIQA, BoolQ, SciQ — generic LM quality, no GSM8K, no MATH, nothing that would catch the 'latent reasoning' benefit the paper keeps invoking as its motivation. 25 00:12:20,100 --> 00:12:44,225 [Hal Turing] On novelty — as we noted, the loop itself traces back to Universal Transformers, so Parcae's real contribution is the stability fix wrapped around Geiping and colleagues' 2025 recurrent-depth paper's RDM baseline, not the looping idea itself. Given that, how do we know truncated-backprop instability is even the right problem to be solving in the first place? 26 00:12:44,225 --> 00:13:28,125 [Dr. Ada Shannon] Because there's a whole competing paradigm the paper rules out by scope rather than comparison. Section B sets aside deep equilibrium models — Bai, Kolter and Koltun's 2019 Deep Equilibrium Models paper — which get constant-memory, infinite-depth training without truncated BPTT. If the motivation is memory-constrained edge deployment, DEQs solve that more directly, and there's no head-to-head against them. The FLOP accounting itself uses McLeish and colleagues' 2025 retrofitted-recurrence convention for 'effective parameters' rather than measured wall-clock GPU time, so 'looping is orthogonal to scaling' hasn't been checked against real hardware cost. And for a paper motivated by edge deployment, there's zero latency, memory, or throughput measurement on actual devices — parameter-count arithmetic standing in for a hardware claim. 27 00:13:28,125 --> 00:14:01,650 [Hal Turing] Practically, this matters most for anyone serving small models under real latency or memory constraints — it gives you a plannable knob across parameters, data, recurrence, and test-time steps instead of pure guesswork. It's also worth cross-checking against Bae and colleagues' 2025 Mixture-of-Recursions, which chases the same adaptive-compute goal but with explicit per-token halting instead of Parcae's implicit depth sampling. Nobody's benchmarked the two against each other yet, which feels like an obvious next paper. 28 00:14:01,650 --> 00:14:23,050 [Dr. Ada Shannon] The paper's own limitations section is refreshingly specific: loop-unit placement and composition are unresolved, behavior at larger FLOP budgets is unverified, and higher training recurrence demands more test-time steps to actually cash in on it — you're paying that compute cost twice, once in training and once at inference. That's the honest scope of a first paper on a stabilized architecture, not a finished scaling law. 29 00:14:23,050 --> 00:14:50,500 [Hal Turing] So the real takeaway — Parcae turns looped models from something that randomly explodes into something you can actually put a scaling law on, and that's a legitimate contribution. But the laws are validated at two small sizes, the headline perplexity and headline benchmark numbers come from two different experimental setups, and the test-time story has a ceiling nobody's pushed past yet. Worth watching, not worth over-trusting. Thanks for listening, everyone — we'll catch you next time.