Average Pass@1 gain while still shortening traces.
Reported length cut relative to the strongest explicit baseline.
Shorter outputs under the short-context recipe.
Pass@1 improvement despite shorter traces.
The paper’s central causal story, drawn as a pressure diagram rather than prose: long-trace reasoning is pushed through a 4K bottleneck, outputs shorten quickly, and the gradient gets noisy when the final verifier is treated as if it perfectly explains each step.
Newline-delimited steps become the unit of filtering. Toggle between rollout types and credit policies to see which steps get reinforced, ignored, or punished.
The claim is not “more intelligence.” The claim is a better spot on the frontier between answer quality, reasoning length, and stability under the same short-context pressure.
The method sits on a timeline: first expose reasoning, then sample more of it, then prune it, then make compute itself tunable. SAS lands in the “credit assignment under compression” pocket.