AI Post Transformers · Episode Companion

Post-Training Science: Scaling Laws for SFT and LoRA

Baseten researchers turn SFT hyperparameters — learning rate, batch size, LoRA rank, epochs, optimizer — into empirical questions instead of inherited folklore, sweeping one variable at a time across Qwen3 and Llama, 0.6B to 235B, dense and mixture-of-experts.

📄 Read the paper (PDF) 🎧 Listen to the episode
0.6B–235B
Parameter range
2
Model families
4
Production datasets
72
Matched LoRA vs Full FT runs

The iSFT Data Loop — and Where It Closes on Itself

Illustrative reconstruction of the iterative-SFT data pipeline described in the episode, not a figure from the paper.

The wrinkle: the same evaluator that grades training drafts also grades the fine-tuned model's final output — one loop judging itself twice.

Scale Coverage: Two Families, Two Architectures

Hover a point for model details. Circle size and diamond size scale with parameter count (log).

Optimal Learning Rate Is Flat Across Scale

LoRA's optimal LR holds at 1e-3 across two orders of magnitude in model size; full fine-tuning's optimum sits roughly 33× lower, also flat. The 80B full fine-tuning cell was initially run off-rule before correction.

Learning-Rate Grid: Does the Flat Rule Hold Up?

Each cell = normalized validation loss for that (model, learning rate) pair — cooler is better. ★ marks the best rate per row. The MoE 30B model (3.3B active) fine-tunes like a dense ~8.5B model, and the flat rule was the discrete-best rate in 13 of 16 tested cells.

Batch Size Is a Cost Knob, Not a Quality Knob

Smaller batches buy slightly lower loss at a fixed token budget; larger batches buy cheaper wall-clock time. There's no single best value — just a frontier.

LoRA Recovers a Median 98% of Full Fine-Tuning's Gain

Full fine-tuning wins all 72 matched comparisons across four production datasets — but usually by a small margin, while training only 3–13% of the parameters.

Rank × Alpha: Where Extra Capacity Stops Paying Off

★ = recommended default (rank 64, alpha 32). Rank 128 roughly doubles trainable parameters for under a thousandth of a nat of extra quality.

Validation Loss Bottoms Out at ~2 Epochs — the Judge Disagrees

Validation loss overfits past epoch 2 (worse by 40–100% on harder tasks by epoch 8), judged task score holds or rises the whole way, and general instruction-following (IFEval) erodes steadily throughout.

Muon vs AdamW — a Narrow, Real Win

Tested on full fine-tuning only, at roughly a 3× lower learning rate than AdamW. Never tested under a LoRA adapter.

References

  1. Post-Training Science: Scaling Laws for SFT and LoRA — Charles O'Neill, Mudith Jayasekara, Harry Partridge (Baseten), 2026. Paper PDF
  2. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen (Microsoft Research), ICLR 2022. Scholar
  3. QLoRA: Efficient Finetuning of Quantized LLMs — Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer (University of Washington), NeurIPS 2023. Scholar
  4. LIMA: Less Is More for Alignment — Chunting Zhou, Pengfei Liu, et al. (Meta AI / collaborators), NeurIPS 2023. Scholar
  5. Muon: An Optimizer for Hidden Layers / Muon Is Scalable for LLM Training — Keller Jordan et al. (2024); Moonshot AI / Kimi team, 2025. Scholar
  6. Model Soups — Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, et al., 2022. Scholar
  7. Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning — Armen Aghajanyan, Luke Zettlemoyer, Sonal Gupta, 2020. Scholar
  8. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng, Shiyi Cao, Dacheng Li, et al., 2023. Scholar