A visual read on agentic test-time scaling: when coding tasks sprawl across tools, files, retries, and failing tests, the paper argues that the real bottleneck is not producing more rollouts, but compressing prior rollouts into summaries that a later model call can actually compare and reuse.
The page starts with the paper’s central move: the useful object is neither the raw transcript nor the final patch, but a structured summary that is small enough to judge and rich enough to carry causal detail.
Instead of one judge staring at all attempts at once, the paper’s parallel mechanism compares small groups recursively. That shrinks the selector’s context and makes pairwise tradeoffs explicit.
Each row is a rollout summary; each column is a memory field. Warmer cells mean the field carries more decisive signal for later selection or refinement under this mock workload.
The second mechanism is not just selection. It distills lessons from prior summaries, then injects those lessons into a later rollout to change what the next attempt actually tries.
These curves are illustrative, not sourced from the paper’s exact tables. They visualize the podcast’s core question: when budgets rise, are gains coming from more attempts, better summary-mediated selection, or genuine reuse across attempts?
Long-horizon coding does not spend its inference budget in one place. The chart below splits mock test-time compute across acting, summarizing, judging, and refinement.
The transcript pushes on whether the gains are a deep scaling insight or just cleaner orchestration. This matrix contrasts where summary-mediated selection helps most, and where reuse or verifier-heavy methods may dominate instead.
The paper sits between agent memory work and test-time scaling work. The through-line is simple: extra compute only matters if previous attempts become usable state rather than dead transcript.