AI Post Transformers • Visual Companion

Trajectory Summaries for Long-Horizon Coding Agents

A visual read on agentic test-time scaling: when coding tasks sprawl across tools, files, retries, and failing tests, the paper argues that the real bottleneck is not producing more rollouts, but compressing prior rollouts into summaries that a later model call can actually compare and reuse.

arXiv: 2604.16529 Theme: represent → select → reuse Mode: long-horizon coding Visuals: tournament, heatmap, scaling curves
parallel summary selection sequential refinement tool-heavy rollouts context compression
Raw trajectory 200+
steps of shell, file edits, tests, reversals
Middle layer 1 note
hypothesis, evidence, touched files, blockers
Selector view 8→2→1
recursive tournament instead of one giant judge prompt
Refinement loop N+1
later attempt conditioned on distilled lessons

Overview: where the summary sits

The page starts with the paper’s central move: the useful object is neither the raw transcript nor the final patch, but a structured summary that is small enough to judge and rich enough to carry causal detail.

Hover nodes for what each stage preserves and discards. interactive: hover + tabbed navigation
Rollout steps are noisy and high-volume: shell output, file reads, failed tests, partial edits.
Trajectory summaries preserve hypothesis, touched files, evidence, partial progress, and blockers.
Decision layers either compare summaries in parallel or feed distilled lessons into the next attempt.
The transcript frames this as a context-management problem: raw traces are too bulky, final answers are too thin, and the summary becomes the scaling interface.
Mock task used here: a messy multi-file bugfix with unstable tests, config edits, and repeated branch reversals.

Tournament bracket over summarized attempts

Instead of one judge staring at all attempts at once, the paper’s parallel mechanism compares small groups recursively. That shrinks the selector’s context and makes pairwise tradeoffs explicit.

Round 1: summaries compete on progress signal, failure severity, and confidence in the causal story.
Round 2: the selector sees fewer but denser candidates, reducing transcript noise.
Winner: selected patch may still be imperfect, but it is chosen from structured notes rather than giant logs.

Summary anatomy heatmap

Each row is a rollout summary; each column is a memory field. Warmer cells mean the field carries more decisive signal for later selection or refinement under this mock workload.

Hover any cell to inspect field relevance. interactive: toggle + heatmap hover
cool = low value for downstream choice
mid = useful context, but not decisive
hot = likely to drive selection or correction
The paper’s claim only works if lossy compression keeps the causal story. In debugging, exact failure evidence stays hot. In refactoring, structural plan and file scope gain weight.

Step-by-step sequential refinement

The second mechanism is not just selection. It distills lessons from prior summaries, then injects those lessons into a later rollout to change what the next attempt actually tries.

Mock scaling behavior under matched compute styles

These curves are illustrative, not sourced from the paper’s exact tables. They visualize the podcast’s core question: when budgets rise, are gains coming from more attempts, better summary-mediated selection, or genuine reuse across attempts?

Pass@1 baseline
31%
Best-of-8 raw
43%
Summary tournament
54%
Refine-after-summaries
59%

Budget allocation bar chart

Long-horizon coding does not spend its inference budget in one place. The chart below splits mock test-time compute across acting, summarizing, judging, and refinement.

Selection vs reuse matrix

The transcript pushes on whether the gains are a deep scaling insight or just cleaner orchestration. This matrix contrasts where summary-mediated selection helps most, and where reuse or verifier-heavy methods may dominate instead.

Selection wins when many attempts exist and transcript noise is high.
Reuse wins when later attempts can exploit distilled failure lessons.
Verifier pressure rises when environments expose strong partial ground truth.
The skeptical reading from the episode remains visible here: if a stronger selector alone explains most gains, the story is better engineering; if conditioned follow-up attempts solve tasks absent from the original pool, reuse looks more fundamental.

Reference frame

The paper sits between agent memory work and test-time scaling work. The through-line is simple: extra compute only matters if previous attempts become usable state rather than dead transcript.

Scaling Test-Time Compute for Agentic Coding
arXiv:2604.16529
Scaling Test-time Compute for LLM Agents
agent-level scaling
Agentic Test-Time Scaling for WebAgents
web-agent analogue
Procedural Memory Retrieval
memory benchmark
Explicit Information Transmission for Context Compression
related podcast episode