The central question is not whether 13 numbers “contain math.” It’s whether reinforcement learning can steer latent capability already present in a pretrained model, using an ultra‑tiny control surface.
13
trainable scalars in the headline setting
26 B
bf16 storage for the adapter
~95%
GSM8K score at ~10k params (reported trend)