1 00:00:01,000 --> 00:00:40,566 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is Learning, Fast and Slow: Towards LLMs That Adapt Continually — Rishabh Tiwari et al., nine co-authors, out of UC Berkeley, Mila, UT Austin, Eragon, Periodic Labs, and Mirendil, posted to arXiv on May 14th, 2026. Ada, right in the abstract there's a number that stopped me cold: they claim up to three times the sample efficiency of plain reinforcement learning, and up to seventy percent less drift from the base model. Three x and seventy percent, same runs, same method. 2 00:00:40,566 --> 00:01:05,319 [Dr. Ada Shannon] Because it's aimed at something that quietly bugs everyone doing RL post-training: update the weights to get better at math or code, and you pay for it elsewhere. The model forgets things, gets rigid, sometimes gets worse at learning whatever comes next. This paper's bet is that burden doesn't all have to sit on the weights — some of it can live somewhere cheaper. That's the game, and why that split even works is what I want to dig into. 3 00:01:05,319 --> 00:01:25,706 [Hal Turing] Before we go further, let's plant a flag — when you say RL post-training here, you specifically mean RLVR, reinforcement learning from verifiable rewards, right? For listeners who haven't hit that term yet, how's it actually different from the RLHF everyone already knows from ChatGPT? 4 00:01:25,706 --> 00:02:01,186 [Dr. Ada Shannon] Right, RLVR. Instead of a human or a learned reward model scoring output — that's RLHF — you've got an automatic verifier: did the math problem land on the right answer, did the code pass its tests. No learned proxy in the loop. It's cheap, hard to game, and the backbone of most reasoning-RL right now — DeepSeek-R1, DeepSeek AI, 2025, is the poster child. But RLVR still moves the model's actual parameters via gradient descent, same as any RL method. That's the 'slow' side here, exactly where those costs start piling up. 5 00:02:01,186 --> 00:02:14,421 [Hal Turing] So that's the slow side, and I get why it's costly. What's the fast side you keep circling back to — the thing that's supposedly cheap enough to sidestep all that, without ever touching a single parameter? Because that sounds almost too easy. 6 00:02:14,421 --> 00:02:52,270 [Dr. Ada Shannon] The fast side is in-context learning — what you and I do any time we tweak a system prompt or add a few examples to a query. Nothing about the parameters changes; you're only changing what sits in the context window before generation happens. It's instant, cheap, and doable per request — Brown et al.'s GPT-3 paper, OpenAI, 2020, showed that alone can shift behavior. The catch: historically it's never matched what weight updates buy you. You optimize a prompt, get a bump, then hit a ceiling weight updates don't. That asymmetry is the whole tension this paper's trying to dissolve. 7 00:02:52,270 --> 00:03:08,710 [Hal Turing] And here's what I don't fully get — you said weight updates cost something. What's the actual bill? 'Catastrophic forgetting' gets thrown around constantly, and I want listeners to have a concrete picture of what breaks — because I have a feeling there's more than one thing going on under that label. 8 00:03:08,710 --> 00:03:28,586 [Dr. Ada Shannon] Two separate line items, and people conflate them. First, catastrophic forgetting: RL a model hard on math, and its ability to write a decent paragraph or hold a normal conversation degrades, because the same weights holding that general competence just got shoved toward maximizing reward on math problems. Second, and this is the subtler one — 9 00:03:28,586 --> 00:03:40,753 [Hal Turing] Wait, sorry, hold on — that's actually the part I've watched happen firsthand. A model gets great at one narrow benchmark and somehow forgets how to just answer a normal question. That's forgetting, full stop? 10 00:03:40,753 --> 00:04:27,286 [Dr. Ada Shannon] Exactly that. The subtler bill is plasticity loss — enough of these updates, and the model doesn't just forget old material, it gets worse at learning anything new. The parameters specialize into a narrow, stiff region where gradients stop being useful. Dohare, Hernandez-Garcia, and colleagues from Richard Sutton's group at the University of Alberta nailed this down in Nature, 2024 — 'Loss of Plasticity in Deep Continual Learning.' They showed ordinary backprop networks can get so stiff after enough sequential tasks that they end up worse than a plain linear model. That's the regime this paper actually cares about — continual learning, absorbing new tasks on the fly without losing the ability to absorb the next. 11 00:04:27,286 --> 00:04:40,568 [Hal Turing] So System 1 and System 2 — that's the analogy they lean on here, fast intuitive judgment versus slow deliberate reasoning? Is that literal, or just a label borrowed because it makes a good story for a training paper to lean on? 12 00:04:40,568 --> 00:05:07,224 [Dr. Ada Shannon] That's their framing. Slow weights play System 2 — deliberate, effortful, expensive to change. Fast weights, the context, play System 1 — quick, cheap, immediately available. Don't push the metaphor too hard; humans aren't running gradient descent on synapses in real time. But it captures the point: you want both timescales working together instead of routing every lesson through the slow, expensive channel. 13 00:05:07,224 --> 00:05:20,878 [Hal Turing] Okay — so given forgetting, plasticity loss, and the ceiling on pure prompting, what's the actual headline number this paper is claiming ties all that together? Because right now it just sounds like a wish list, not a result. 14 00:05:20,878 --> 00:06:12,797 [Dr. Ada Shannon] Figure 1 is basically the whole paper in one picture. Combine both channels — what they call Fast-Slow Training — and it reaches RL's peak accuracy using up to three times fewer training samples, converging to a higher ceiling than either RL or prompting alone. It also drifts up to seventy percent less from the base model, by KL divergence, their forgetting proxy. And because it drifts less, it stays plastic — train it on one task, and it can still pick up a second, where a pure-RL model basically stalls. Every one of those claims deserves scrutiny, and that's exactly what the rest of this episode is for. And that's the outcome, not the mechanism — this isn't RL first with a good prompt bolted onto the finished checkpoint afterward. The two channels run inside the same loop, trading off which one absorbs the adaptation as training goes. 15 00:06:12,797 --> 00:06:19,113 [Hal Turing] Okay, make that concrete for me. When you say they run in the same loop, what's literally updating, and on what schedule? 16 00:06:19,113 --> 00:07:20,553 [Dr. Ada Shannon] The slow side is CISPO, the surrogate loss from the ScaleRL recipe — updating theta, the model weights, off scalar reward alone, same as ordinary RLVR. The fast side is GEPA, a reflective evolutionary optimizer, and instead of touching parameters it rewrites a population of K prompts using the full rollout text — the reasoning trace, tool calls, and error messages a scalar reward throws away. Training runs in cycles of six RL steps. At the top of each cycle, GEPA looks at the current policy, evolves the prompt population using that richer text signal, and hands back K fresh candidates. RL then trains on those prompts for six steps before the next GEPA cycle fires. It's a population, not a single winning prompt, because GEPA naturally returns a Pareto frontier — different prompts specialize on different slices of the problem set, so sampling across them gives RL's group-relative advantage richer variation than just resampling one prompt. 17 00:07:20,553 --> 00:07:27,798 [Hal Turing] And does splitting the work across two channels actually buy something concrete, or is this a clean idea that nets out even with RL alone? 18 00:07:27,798 --> 00:08:07,736 [Dr. Ada Shannon] Very concrete. On raw data efficiency, FST matches RL's peak validation accuracy in three times fewer training steps on CodeIO, one-point-four times fewer on Polaris math, three times fewer on HoVer-hard — and it costs nothing out of distribution, the OOD averages land essentially flat between the two. But it doesn't just get there faster, it goes further. They fit a saturation curve to each run and read off where it's actually converging, not wherever training happened to stop, and FST's asymptote sits higher on all three: plus four-point-four points on CodeIO, plus two-point-nine on Polaris, plus seven-point-seven on HoVer-hard. 19 00:08:07,736 --> 00:08:17,442 [Hal Turing] Wait, hold on — seven points higher on HoVer-hard isn't rounding error, that's a genuinely different ceiling out of the same base model and the same reward signal. 20 00:08:17,442 --> 00:09:12,381 [Dr. Ada Shannon] Right, and it connects straight to why the model drifts less. At matched accuracy, FST sits at meaningfully lower KL to the base policy than RL on CodeIO, HoVer, and Physics — that's the seventy percent number from the headline. And it's not a strawman comparison: Shenfeld and colleagues out of MIT showed in 2025, in the RL's Razor paper, that on-policy RL is already biased toward KL-minimal solutions relative to offline methods — that's already the strong baseline. FST's claim is it shifts that frontier further left still. Where it gets visceral is the plasticity probe: train on Polaris math, then run fresh RL on HoVer-hard — the RL-trained checkpoint collapses HoVer-hard learnability to basically zero. FST-init stays close to a from-scratch reference. Same story from Physics: FST hits twenty-four-point-two percent on HoVer-hard at step four hundred, RL-init is stuck at nineteen-point-nine. 21 00:09:12,381 --> 00:09:18,139 [Hal Turing] So in a continual setting, where the tasks never stop arriving, does that gap actually compound? 22 00:09:18,139 --> 00:09:48,836 [Dr. Ada Shannon] That's Advantage five, and it's the sharpest result in the paper. One uninterrupted training run that swaps the task every two hundred steps — HoVer, then CodeIO, then Physics. FST reaches near-peak accuracy at every stage. RL picks up HoVer fine in stage one, then hits CodeIO in stage two and stalls — barely lifts off its starting accuracy for the entire two-hundred-step budget. Whatever RL spent learning HoVer left it specialized enough that CodeIO barely moves the needle at all. 23 00:09:48,836 --> 00:09:58,124 [Hal Turing] So why does the fast channel actually get there sooner, mechanically — what's it doing that gradient descent on theta can't do in the same number of steps? 24 00:09:58,124 --> 00:10:50,276 [Dr. Ada Shannon] They isolate that with a synthetic star-graph search task where the base model starts at essentially zero reward. Parameter-only RL sits near zero for roughly the first two-fifty to three-hundred steps before it moves at all. FST breaks out of zero by around step fifty — almost an order of magnitude sooner — because GEPA reads the actual failure text, exactly where a path went wrong, and injects that structure into the prompt population before theta has moved appreciably. The gain decomposition tells a more mixed story, though: on HoVer-hard and CodeIO, both channels contribute and combining them clearly beats either alone. On Polaris math, almost all the gain is carried by the slow weights, twenty percent up to forty-seven, with the fast channel barely moving the needle. That's the one task where this whole story looks different, and it's worth digging into why. 25 00:10:50,276 --> 00:10:55,524 [Hal Turing] Before we get there — none of this is free. Two optimization loops running at once has to cost something. 26 00:10:55,524 --> 00:11:22,877 [Dr. Ada Shannon] It does. Per RL step, FST runs about a hundred seconds against roughly sixty for RL alone, because GEPA's rollouts and its reflection-model calls are real wall-clock sitting on top of the RL loop. A rollout-reuse trick brings it down to about forty-seven seconds by recycling GEPA's evaluation rollouts back into the RL group instead of generating everything fresh, but that's a mitigation layered on afterward, not the baseline the headline efficiency numbers assume. 27 00:11:22,877 --> 00:11:59,239 [Hal Turing] Right, forty-seven seconds — but that's just the RL loop. Every cycle, GEPA still runs its own rollouts across the K candidate prompts and pays for calls to the reflection model, and none of that shows up in the per-step number. The paper admits it straight out in an appendix — end to end, a full FST run costs more than an RL-only run of the same step count. So when the headline says one-point-four to three times fewer training steps, is that a real compute or dollar win, or a step-count win hiding the actual bill? 28 00:11:59,239 --> 00:12:40,153 [Dr. Ada Shannon] Mostly the latter. A headline run burns twenty-five to forty H100-GPU-hours, and GEPA eats a real chunk of it — the reflection calls alone go through gpt-5.2 over LiteLLM, running around ten dollars a run, billed separately from the compute. Honestly, what's actually new here isn't the RL half — that's Khatri, Madaan, and Agarwal's own ScaleRL and CISPO recipe from 2025, largely the same UT Austin and Mila group, wearing a different hat. The genuinely new piece is the interleaved GEPA channel bolted onto it. Which also means the RL baseline is this group's own already-tuned recipe, not an independent one. 29 00:12:40,153 --> 00:13:01,469 [Hal Turing] Oh — wait, hold on, that Khatri connection is actually what I wanted to ask about, because it ties right into something in Appendix G. For Polaris, for math, the RL and FST curves in KL-versus-reward space just sit right on top of each other. No leftward shift. And Figure 5 in the main text quietly drops the Polaris panel entirely. 30 00:13:01,469 --> 00:13:41,778 [Dr. Ada Shannon] Yeah, that's a real tell. Every other task trains on Qwen3-8B-Instruct, already strong at following formatting and self-checking instructions. Polaris can't use that checkpoint — the public model's already saturated on math — so they built a custom base, SFT'ing Qwen3-8B-Base on Nemotron data. That base is measurably worse at instruction-following, and GEPA's whole mechanism depends on a policy that actually reads and obeys an evolved prompt. Take that away and the fast channel has nothing to push against. Reward still climbs from RL alone, but the KL story collapses — so the plasticity benefit may be a property of starting from a model already instruction-tuned enough to listen. 31 00:13:41,778 --> 00:14:21,485 [Hal Turing] And that's the dataset in miniature. Every headline number comes from Qwen3-8B, plus one Qwen3-4B-Instruct run on the synthetic star-graph task. No seventy-billion-parameter run, nothing checking whether this is a property of fast-slow optimization generally or just how receptive this one heavily RLHF'd checkpoint is to prompt steering. And none of the headline comparisons — the sigmoid fits, the KL frontier, the plasticity probe — carry error bars or multiple seeds. Nineteen-point-nine versus twenty-four-point-two on the Physics-to-HoVer probe gets presented as a real gap, and RL training is famously noisy run to run. 32 00:14:21,485 --> 00:14:53,760 [Dr. Ada Shannon] Fair worry, and it connects to something the paper leans on without measuring. The plasticity narrative traces back to that Alberta paper we mentioned earlier — the one that mechanistically defines plasticity loss: growing weight and activation norms, effective-rank collapse, literal dead units. FST never measures any of those directly. It substitutes KL-to-base and Phase-2 accuracy as proxies. Reasonable stand-ins, but not the same thing — a model can sit at low KL to base and still have collapsed-rank layers underneath, and these plots would never show it. 33 00:14:53,760 --> 00:15:17,956 [Hal Turing] And the KL story has a ceiling on how novel it is. RL's Razor already showed that the size of that KL shift predicts how much gets forgotten. FST's pitch is that it pushes the frontier further left than that already-KL-minimal baseline. So how much of the seventy-percent number is genuinely new, versus riding an effect Shenfeld already documented, plus whatever the prompt channel adds? 34 00:15:17,956 --> 00:15:55,340 [Dr. Ada Shannon] Some of both, and the paper doesn't let you cleanly separate it. Where I'd deploy this is narrow — continual settings, a model facing a real stream of shifting tasks that can't afford to go stale, which is exactly the HoVer-to-CodeIO-to-Physics run they built. For one-off fine-tuning on a single task, GEPA's overhead probably isn't worth it over plain RL. And they're upfront about a ceiling on the fast-only side too — Appendix H tries folding the prompt's gains into the weights through reverse-KL distillation, no reward involved, and it plateaus well below full FST. Both channels still need to chase reward jointly. 35 00:15:55,340 --> 00:16:25,200 [Hal Turing] So that's the shape of it — a real, replicated data-efficiency and ceiling win on three of four tasks, a KL-reduction claim that quietly doesn't survive its own appendix on math, and a plasticity story resting on proxies for something the paper never directly measures. Genuinely useful direction, an oversold headline number. That's 'Learning, Fast and Slow.' Thanks for listening — take care.