A controller learns how to spend inference-time compute across branching, continuation, probing, pruning, and stopping. The page below focuses on the geometry of that policy search: where budget moves, how replay makes search cheap, and why the real object being optimized is a compute allocation strategy rather than more tokens.
2605.08083
The paper’s key move is to convert test-time scaling into a controller search problem. A small action grammar manages breadth, depth, verification, and budget while offline replay scores candidate policies against logged traces.
Three ledgers matter: data collection, cheap offline search, and inference-time execution of the discovered controller.
This controller surface shows which action dominates for a given state. Hover any cell to inspect the preferred move and expected payoff under a compact one-parameter tradeoff family.
Budget is the vertical axis, branch depth is horizontal, and color intensity reflects estimated controller confidence.
The discovered controller is interesting only if it improves the frontier, not just raw accuracy. Toggle benchmark families to see how different policies trade tokens, latency, and branch explosion.
Frontier gains come from spending budget where branches still have upside and exiting quickly once probes become decisive.
Logged traces make policy evaluation cheap because the controller can be replayed over previously collected trajectories. Step through a synthetic problem and watch how the same trace supports multiple policy choices.
Trace feedback is fine-grained: the search loop can see not just win or loss, but exactly where a candidate policy overspent or pruned too early.
The companion centers on controller synthesis, offline replay, and test-time compute allocation. References below emphasize the source paper, control-theoretic ancestry, offline RL, and adjacent podcast episodes.