AI Post Transformers arXiv:2605.08083 2026 episode companion

Agentic Discovery for Test-Time Scaling

A controller learns how to spend inference-time compute across branching, continuation, probing, pruning, and stopping. The page below focuses on the geometry of that policy search: where budget moves, how replay makes search cheap, and why the real object being optimized is a compute allocation strategy rather than more tokens.

Controller verbs

Action 1
branch
Action 2
probe
Action 3
prune

Search economics

Logged traces
128
Probe interval
500t
Quoted discovery
$39.9

Transcript arXiv IDs

2605.08083

Discovery loop as a visible control system

The paper’s key move is to convert test-time scaling into a controller search problem. A small action grammar manages breadth, depth, verification, and budget while offline replay scores candidate policies against logged traces.

logged data search loop live inference

Current lens

Three ledgers matter: data collection, cheap offline search, and inference-time execution of the discovered controller.

3 ledgers

State vector

Policy heatmap over branch depth and budget

This controller surface shows which action dominates for a given state. Hover any cell to inspect the preferred move and expected payoff under a compact one-parameter tradeoff family.

continue branch probe prune stop

Hovered state

Budget is the vertical axis, branch depth is horizontal, and color intensity reflects estimated controller confidence.

branch

Action mix by regime

Accuracy-cost frontier

The discovered controller is interesting only if it improves the frontier, not just raw accuracy. Toggle benchmark families to see how different policies trade tokens, latency, and branch explosion.

Auto-discovered policy manual tree search fixed sampling

Selected operating point

Frontier gains come from spending budget where branches still have upside and exiting quickly once probes become decisive.

+6.3 acc/token-k

Budget composition

Offline replay inspector

Logged traces make policy evaluation cheap because the controller can be replayed over previously collected trajectories. Step through a synthetic problem and watch how the same trace supports multiple policy choices.

surviving branch inactive trace pruned path

Step feedback

Trace feedback is fine-grained: the search loop can see not just win or loss, but exactly where a candidate policy overspent or pruned too early.

probe opens

Support mismatch meter

References

The companion centers on controller synthesis, offline replay, and test-time compute allocation. References below emphasize the source paper, control-theoretic ancestry, offline RL, and adjacent podcast episodes.