AI Post Transformers · Visual Companion

Benchmarking Test-Time Scaling for General LLM Agents

Paper 2602.18998 arXiv ↗ Mixed benchmark tasks 496 Agent families evaluated 10

This page visualizes the paper’s central move: take search, coding, reasoning, and API/tool-use tasks from multiple benchmarks, strip away benchmark-specific scaffolding, and force one agent to infer task type, select tools, and plan under one shared interface.

~30%
Typical performance drop in general vs. standard setting
2 modes
Sequential scaling vs. parallel scaling
2 bottlenecks
Context overload and self-selection failure
Search / Browse Coding / Terminal Reasoning / Tools

From specialized sandboxes to one general interface

The benchmark’s key intervention is the wrapper: same prompt style, same interaction loop, broad shared tool pool, and no benchmark-specific cues telling the agent what kind of task it is in.
BrowseComp 124 WebVoyager 65 SWE-Bench 50 Terminal-Bench 80 MathHay 75 Tau2-Bench 50 MCP-Bench 52
What changes? Not the underlying task sources, but the interface ergonomics and the need to infer intent, route tools, and adapt online.
Why this matters Standard benchmark setups often pre-solve parts of the problem by labeling the domain and narrowing the action space.
Visual cue Toggle between modes to see how environment specialization collapses into one neutral, more ambiguous field.

Cross-benchmark degradation under a shared wrapper

Mocked but paper-faithful numbers illustrate the reported pattern: broad drops in the general setting, with search and tool-use especially sensitive to interface unification. Hover cells for values.
Reported headline Most agents lose roughly thirty percent of performance when benchmark-specific scaffolds are removed.
Asymmetric damage Tool calling and web tasks degrade more than coding on some systems, suggesting wrapper sensitivity and tool arbitration errors.
Family effect The transcript notes Claude-family models as comparatively robust under the unified setting.

Sequential vs. parallel test-time scaling

The paper separates two ways to spend extra inference compute: more turns in one trajectory, or more independent trajectories. Their shapes differ: context clutter hurts long traces, while best-of-K is capped by weak self-selection.
Sequential result Gains rise early, then flatten or decline as transcript length, stale plans, and tool traces become hard to use productively.
Parallel result pass@K keeps climbing, meaning good candidates exist somewhere in the sample set.
But Self-choice stays flat, so practical utility depends on verification and selection, not just generation diversity.

What is actually failing? Decision-loop decomposition

A general benchmark is most useful when it separates failure sources: intent inference, tool choice, interface syntax, memory over time, and answer verification. This section visualizes that decomposition.
Research takeaway One aggregate score hides whether an agent misunderstood the request, chose the wrong tool, or generated a correct answer and failed to recognize it.
Practical takeaway Product teams should evaluate routing, tool arbitration, state compression, and checking—not just end success.
Open question Some degradation may reflect true robustness limits; some may be wrapper unfairness. Better benchmarks disentangle both.

References

Compact source list used for this visual companion. arXiv IDs found in provided material: 2602.18998.
Li et al., 2026
Benchmark Test-Time Scaling of General LLM Agents
arXiv:2602.18998
Jimenez et al., 2023
SWE-Bench
Scholar link
Aleithan et al., 2024
Terminal-Bench
Scholar link
Zhou et al., 2023
WebVoyager
Scholar link
BrowseComp
Benchmark source for browsing/search tasks
Scholar link
Related TTS work
Self-consistency, verifier training, Quiet-STaR, Snell et al.
Scholar link