The iSFT Data Loop — and Where It Closes on Itself
Illustrative reconstruction of the iterative-SFT data pipeline described in the episode, not a figure from the paper.
Scale Coverage: Two Families, Two Architectures
Hover a point for model details. Circle size and diamond size scale with parameter count (log).
Optimal Learning Rate Is Flat Across Scale
LoRA's optimal LR holds at 1e-3 across two orders of magnitude in model size; full fine-tuning's optimum sits roughly 33× lower, also flat. The 80B full fine-tuning cell was initially run off-rule before correction.
Learning-Rate Grid: Does the Flat Rule Hold Up?
Each cell = normalized validation loss for that (model, learning rate) pair — cooler is better. ★ marks the best rate per row. The MoE 30B model (3.3B active) fine-tunes like a dense ~8.5B model, and the flat rule was the discrete-best rate in 13 of 16 tested cells.
Batch Size Is a Cost Knob, Not a Quality Knob
Smaller batches buy slightly lower loss at a fixed token budget; larger batches buy cheaper wall-clock time. There's no single best value — just a frontier.
LoRA Recovers a Median 98% of Full Fine-Tuning's Gain
Full fine-tuning wins all 72 matched comparisons across four production datasets — but usually by a small margin, while training only 3–13% of the parameters.
Rank × Alpha: Where Extra Capacity Stops Paying Off
★ = recommended default (rank 64, alpha 32). Rank 128 roughly doubles trainable parameters for under a thousandth of a nat of extra quality.
Validation Loss Bottoms Out at ~2 Epochs — the Judge Disagrees
Validation loss overfits past epoch 2 (worse by 40–100% on harder tasks by epoch 8), judged task score holds or rises the whole way, and general instruction-following (IFEval) erodes steadily throughout.
Muon vs AdamW — a Narrow, Real Win
Tested on full fine-tuning only, at roughly a 3× lower learning rate than AdamW. Never tested under a LoRA adapter.