1 00:00:01,000 --> 00:00:31,300 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into 'Post-Training Science for Supervised Fine-Tuning' — Charles O'Neill et al., three authors total, with Mudith Jayasekara and Harry Partridge, all out of Baseten. And I like the ambition here: instead of another fine-tuning trick, they're trying to replace the ad hoc guesswork around SFT with actual controlled measurement. 2 00:00:31,300 --> 00:00:50,950 [Dr. Ada Shannon] What got me is that most fine-tuning papers are 'here's a trick, trust us.' This one is closer to a physics lab notebook — one variable at a time, across real production data, asking whether learning rate, batch size, LoRA rank, epochs, and optimizer choice actually follow rules that transfer, instead of being rediscovered from scratch every time. 3 00:00:50,950 --> 00:01:14,000 [Hal Turing] Which matters because SFT is basically the default last mile for every production model now — Llama, Qwen, whatever you're deploying — and yet those five decisions are usually inherited from pretraining or copied off whatever worked last time. So Ada, ground us: mechanically, what is SFT, and what's LoRA doing differently from just fine-tuning the whole model? 4 00:01:14,000 --> 00:01:57,349 [Dr. Ada Shannon] Same machinery as pretraining — forward pass, cross-entropy loss, backprop — just on a much smaller, curated set of prompt-response pairs instead of trillions of raw tokens. Full fine-tuning updates every weight, which for AdamW means storing two extra copies of every parameter, so at seventy billion-plus parameters that's brutal on hardware. LoRA — Low-Rank Adaptation, from Edward Hu and colleagues at Microsoft Research in 2021 — freezes the whole pretrained model and bolts on small trainable adapter matrices next to the attention and MLP weights. Instead of a full-rank update, you learn a rank-r approximation, r typically eight to sixty-four versus thousands of real dimensions. A hundred to a thousand times fewer trainable parameters, at the cost of capping how much the model can actually change. 5 00:01:57,349 --> 00:02:24,074 [Hal Turing] And they push this all the way to two-thirty-five billion parameters on the mixture-of-experts side, plus a second model family, Llama, layered on top of Qwen3, specifically to see whether anything they find is family-specific or actually general. Give me a quick primer on mixture-of-experts for anyone who hasn't hit that term before, and then I want to get into where their training data actually comes from, because that part's the wildest piece of this whole setup. 6 00:02:24,074 --> 00:02:57,125 [Dr. Ada Shannon] Dense models run every token through every parameter. MoE splits each layer into many expert sub-networks and routes each token to just a few — so two-thirty-five billion total parameters, but only twenty-two billion active per token. As for the data — all four datasets, security, leasing, support, docs, come from iterative SFT: a model drafts an answer, a customer-built evaluator grades it and explains the failure, and it's revised until it passes. So every kept example already satisfies the evaluator, which means no label noise, but it also means the judge scoring quality later is— 7 00:02:57,125 --> 00:03:18,599 [Hal Turing] Wait, wait — hold on. The same judge? The evaluator that decided the training data was good enough is the exact same one they use afterward to grade whether the fine-tuned model's outputs are any good? That's not two independent checks, that's one loop closing back on itself, and that feels like something worth sitting with before we go any further. 8 00:03:18,599 --> 00:03:52,800 [Dr. Ada Shannon] Exactly, and we'll come back to how much weight that can bear. For now: internally consistent targets, scored by the same criterion they were built to satisfy — a controlled testbed and a closed loop at once. Everything they measure gets judged against validation loss, negative log-likelihood on held-out data, and they're borrowing neural scaling laws — the power-law relationship between size and loss from pretraining — to ask whether post-training obeys the same physics, across a one-lever-at-a-time sweep spanning Qwen3, Llama, dense, MoE, LoRA, and full fine-tuning. 9 00:03:52,800 --> 00:04:19,300 [Dr. Ada Shannon] Concretely: the learning-rate result is clean. Across zero-point-six to thirty-two billion parameters, both Qwen3 and Llama, the optimal LoRA rate is flat at ten to the minus three — no drift with scale or family. That's roughly thirty-three times the full fine-tuning optimum, near three times ten to the minus five. Batch isn't a quality knob, it's a cost knob: smaller batch means more steps and slightly lower loss at a fixed token budget; bigger batch is cheaper wall-clock. No single best value, just a frontier. 10 00:04:19,300 --> 00:04:37,074 [Hal Turing] Flat across two orders of magnitude in parameter count is a bold claim to hang on nine models, Ada. Does that rule actually survive contact with something structurally weirder — the mixture-of-experts side they push all the way to two-thirty-five billion parameters? 11 00:04:37,074 --> 00:05:16,050 [Dr. Ada Shannon] It does — that's the surprising part. Held out on a thirty-billion MoE with only three-point-three billion active, the flat rule is the discrete-best learning rate in thirteen of sixteen cells; that model fine-tunes like a dense eight-point-five billion model. They push further, blind, to an eighty-billion and a two-thirty-five-billion MoE, and the size trend still holds — though the full fine-tuning runs there were initially mistuned, three times off the rule, before correction. On LoRA versus full fine-tuning: full fine-tuning wins all seventy-two matched comparisons, but by small margins, and LoRA recovers a median ninety-eight percent of the improvement training just three to thirteen percent of the parameters. 12 00:05:16,050 --> 00:05:31,125 [Hal Turing] Ninety-eight percent of the full fine-tuning gains for a fraction of the parameters is a genuinely strong pitch for LoRA. So does cranking the rank higher just keep buying you more of that missing two percent, or does it run out of room? 13 00:05:31,125 --> 00:06:08,675 [Dr. Ada Shannon] It runs out fast. Rank adds capacity up to about sixty-four then plateaus — rank one-two-eight barely helps, under a thousandth of a nat, while roughly doubling trainable parameters. Alpha thirty-two is best in every cell at rank sixty-four, so the default stays rank sixty-four, alpha thirty-two, with rank thirty-two as the cheap alternative. All of that rests on trusting validation loss, though. Within a fixed model, dataset, and recipe, loss does rank judged quality well — Spearman negative zero-point-three-eight to negative zero-point-eight-eight. But pool across model families and that signal collapses; a lower-loss model can score worse with the judge. 14 00:06:08,675 --> 00:06:15,525 [Hal Turing] Wait, hold on — how does a lower-loss model score worse? Isn't that the whole point of measuring loss? 15 00:06:15,525 --> 00:07:00,125 [Dr. Ada Shannon] Because loss measures fit to text, not the behavior the judge grades. Llama hits lower loss than Qwen at matched size on some datasets while scoring lower with the judges — fit and behavior split once you cross families. That's partly why they also track the Fisher trace, the expected squared gradient norm, a proxy for how flat the loss minimum is, motivated by flatter minima generalizing better at identical loss. But at matched loss within a fine-tuned cell it adds nothing reliable — the sign isn't even stable. Zooming out, loss falls as a clean saturating power law in model size, steeper for full fine-tuning than LoRA, so that ninety-eight percent retention narrows with scale — and all three MoEs land at the geometric mean of their active and total parameters. Tripling fresh examples always lowers loss, though the gain depends far more on dataset than model size. 16 00:07:00,125 --> 00:07:11,575 [Hal Turing] And swapping in Muon for AdamW — does that actually earn its keep here, or is it another pretraining-scale trick that quietly doesn't survive the jump to fine-tuning? 17 00:07:11,575 --> 00:07:32,225 [Dr. Ada Shannon] A narrow keep. Tested on full fine-tuning only, Muon reaches marginally lower loss and a flatter minimum than AdamW on every dataset, but at a learning rate about three times lower — you can't reuse AdamW's rate. On the task judge the two land level. On IFEval, general instruction-following, Muon comes out ahead on every dataset — a retention win, not a task-quality win. 18 00:07:32,225 --> 00:07:42,350 [Hal Turing] So loss overfits, the judge doesn't always agree with it — what's the actual stopping rule? How many epochs before you're just burning capability for nothing? 19 00:07:42,350 --> 00:08:04,525 [Dr. Ada Shannon] About two. Validation loss bottoms out around two epochs, then climbs steadily through eight — forty to a hundred-twenty percent worse on harder tasks. But judged task score doesn't follow; it holds or rises all the way to eight. What erodes the whole time is IFEval, general instruction-following. And giving models fresh data instead of repeating the same five thousand examples won every time at a matched budget — repetition wasn't worth it. 20 00:08:04,525 --> 00:08:53,100 [Hal Turing] Ada, here's the thing that's been nagging me since you laid out that judge overlap in Part 1. Every one of these datasets comes from drafting against an evaluator, revising until it passes, then that same evaluator later grades the fine-tuned model's output as the 'judged quality' score you've cited all episode. So when the paper reports those Spearman correlations — negative point-three-eight to negative point-eight-eight between loss and judged quality within a recipe — how much of that is really telling us something general about validation loss, versus the judge just recognizing its own fingerprints? If you trained on organically authored data with independent gold labels instead of iSFT output, does that relationship survive at all, or does it just measure a closed loop being consistent with itself? 21 00:08:53,100 --> 00:09:35,425 [Dr. Ada Shannon] That's exactly the caveat the authors own — they call the judge 'not independent of the training objective' and lean on IFEval as the one outside check the iSFT loop never touched. But that doesn't fully answer it, because the inside judge's own reliability is thin: it's GPT-5.5 grading four anonymized, single-customer datasets, and the only validation reported is that the rubric was customer-approved — no human-versus-judge agreement rate anywhere. That matters, because the epoch stopping rule, the Muon recommendation, and the rank-sixty-four default are all read off judged score, not loss. Swap GPT-5.5 for a different judge and you're not second-guessing one number — you could reorder which epoch checkpoint looks best, whether Muon's IFEval edge holds, maybe even the rank plateau. 22 00:09:35,425 --> 00:10:19,250 [Hal Turing] Oh wait, hold on — that actually connects to something that bothered me even more: the headline claim that the flat ten-to-the-minus-three LoRA rate 'transfers unchanged to two-thirty-five-billion.' That's not an independent sweep at that scale — it's one blind, untuned cell per dataset. And on the full-fine-tuning side, the eighty-billion cells were initially mistuned, two-point-six times above the rule's predicted rate, before getting moved onto the on-rule value where they matched LoRA again. So 'the rule transfers to two-thirty-five-billion' might really mean 'a single untested guess didn't blow up,' not 'we verified ten-to-the-minus-three is optimal at that scale.' Those are different claims dressed up as the same sentence. 23 00:10:19,250 --> 00:11:03,875 [Dr. Ada Shannon] That's the sharpest gap between what's tested and what's implied in this paper. Which ties into something that surprised me on the references page — no citation to Hu and colleagues' original LoRA paper, out of Microsoft, twenty-twenty-one. The second half of this study is basically a hyperparameter investigation of LoRA, and they cite Kalajdzievski's rank-stabilization work and Biderman's 'LoRA learns less and forgets less,' but not the method paper itself. Two more would've helped. QLoRA — Dettmers and colleagues, University of Washington, twenty-twenty-three — is directly relevant since they flag quantization-aware training as future work, and nobody's tested whether rank-sixty-four, alpha-thirty-two holds once the base model is quantized. And Model Soups — Wortsman and colleagues, University of Washington and Google, twenty-twenty-two — since their own flatness numbers say Muon makes more mergeable checkpoints, and they never connect that to weight averaging. 24 00:11:03,875 --> 00:11:50,925 [Hal Turing] Give them credit, though — the honesty about how thin that scaffolding is happens to be rarer than the scaffolding itself in this literature. Four datasets, six model sizes per family at most, single-customer proprietary data throughout, and the epoch and scaling ladders stop before thirty-two billion because full-context evaluation blows past a single device's memory. The optimizer comparison is full-fine-tuning only, too — Muon never gets tested under a LoRA adapter, which is exactly what most people reading this would deploy. Small tangent, but O'Neill's got another paper out right now on whether language models can learn facts continually in their weights — same pattern, narrow careful instruments over one flashy claim. 25 00:11:50,925 --> 00:12:06,925 [Dr. Ada Shannon] Thirty-two billion, yeah — that's a memory-limit wall, not a hidden result, so I'll give them that one. But it's the right note to end the critique on, because the interesting question now isn't whether this paper is perfect. It's what you actually do with it on Monday morning. 26 00:12:06,925 --> 00:12:28,100 [Hal Turing] So give me the checklist, Ada. Say I'm a team standing up an SFT job tonight on some open base model — Qwen, Llama, doesn't matter which family — and I don't want to burn a week rediscovering hyperparameters this paper already measured. Rank, alpha, learning rate, epochs — walk me through the actual numbers, not just the vibes. 27 00:12:28,100 --> 00:13:10,950 [Dr. Ada Shannon] Start from r equals sixty-four, alpha thirty-two, learning rate ten to the minus three — that combination held up across both families and every scale they tested, so it's a legitimate default, not a guess. Train roughly two epochs, then stop trusting the loss curve and start adding fresh examples instead of looping back over the same data — their own overfitting result says repeats buy you nothing past that point. One operational caveat: only trust validation loss to rank checkpoints inside one fixed model-and-recipe cell. Cross model families or judges and loss stops being your umpire — go straight to the task evaluation. And if you're running full fine-tuning and care about retained instruction-following, swapping AdamW for Muon is close to a free win — nobody's tested it under a LoRA adapter yet, so don't assume it carries over. 28 00:13:10,950 --> 00:13:36,525 [Hal Turing] Oh wait, actually — that untested-under-LoRA thing is exactly what jumped out at me on their own future-work list. Baseten's audience is deploying this stuff, and a lot of production LoRA serving today quantizes the frozen base to save memory. So the rank-sixty-four, alpha-thirty-two default nobody's checked against a quantized base is a real operational gap, not just an academic footnote. 29 00:13:36,525 --> 00:14:13,225 [Dr. Ada Shannon] Right, and quantization-aware training is exactly where they point next — that's the one gap I'd want closed before shipping the default blind. Beyond that: fine-tuning reasoning models on non-reasoning or off-policy traces, how any of this interacts with continued pretraining on a model's own past assistant outputs, and pushing the adapter study past LoRA into learned weight deltas more broadly. They also gesture at rejection-sampling fine-tuning and on-policy RL — really asking whether hand-curating iSFT data offline is just a cheaper stand-in for what on-policy methods do automatically during training. 30 00:14:13,225 --> 00:14:40,400 [Hal Turing] Good note to land on. The real contribution here isn't any single number — it's that somebody measured this instead of copying last quarter's config, one lever at a time, across scale and architecture. Just hold onto the caveat from earlier: that measurement is only as honest as the judge grading it, and here the judge graded its own homework. Controlled experiments on a closed loop are still useful — just don't mistake internally consistent for externally validated. 31 00:14:40,400 --> 00:14:51,226 [Dr. Ada Shannon] Agreed — that's the real discipline this paper models, more than any single number in it: replace folklore with an actual instrument, then be honest about what it can't see. 32 00:14:51,226 --> 00:15:00,576 [Hal Turing] Well said, Ada — good place to leave it. Thanks for going deep on this one with me, and thanks to everyone listening in. We'll catch you next time.