Learning to Reason
with 13 Parameters

A visual companion to the podcast episode on TinyLoRA + RL: can a frozen 7B–8B instruction model get large math gains when only a microscopic set of parameters is updated?
arXiv:2602.04118 Models: Qwen2.5‑7B / 8B Instruct Method: GRPO + TinyLoRA Benchmarks: GSM8K, AIME, AMC, MATH500 Key headline: 91% pass@1 with 13 params
The central question is not whether 13 numbers “contain math.” It’s whether reinforcement learning can steer latent capability already present in a pretrained model, using an ultra‑tiny control surface.
13
trainable scalars in the headline setting
26 B
bf16 storage for the adapter
~95%
GSM8K score at ~10k params (reported trend)
RL/GRPO remains strong deep into tiny budgets
SFT degrades sharply under tiny budgets

References & source trail