Benchmarking Test-Time Scaling for General LLM Agents
Paper 2602.18998arXiv ↗Mixed benchmark tasks 496Agent families evaluated 10
This page visualizes the paper’s central move: take search, coding, reasoning, and API/tool-use tasks from multiple benchmarks, strip away benchmark-specific scaffolding, and force one agent to infer task type, select tools, and plan under one shared interface.
~30%
Typical performance drop in general vs. standard setting
From specialized sandboxes to one general interface
The benchmark’s key intervention is the wrapper: same prompt style, same interaction loop, broad shared tool pool, and no benchmark-specific cues telling the agent what kind of task it is in.
What changes? Not the underlying task sources, but the interface ergonomics and the need to infer intent, route tools, and adapt online.
Why this matters Standard benchmark setups often pre-solve parts of the problem by labeling the domain and narrowing the action space.
Visual cue Toggle between modes to see how environment specialization collapses into one neutral, more ambiguous field.
Cross-benchmark degradation under a shared wrapper
Mocked but paper-faithful numbers illustrate the reported pattern: broad drops in the general setting, with search and tool-use especially sensitive to interface unification. Hover cells for values.
Reported headline Most agents lose roughly thirty percent of performance when benchmark-specific scaffolds are removed.
Asymmetric damage Tool calling and web tasks degrade more than coding on some systems, suggesting wrapper sensitivity and tool arbitration errors.
Family effect The transcript notes Claude-family models as comparatively robust under the unified setting.
Sequential vs. parallel test-time scaling
The paper separates two ways to spend extra inference compute: more turns in one trajectory, or more independent trajectories. Their shapes differ: context clutter hurts long traces, while best-of-K is capped by weak self-selection.
Sequential result Gains rise early, then flatten or decline as transcript length, stale plans, and tool traces become hard to use productively.
Parallel resultpass@K keeps climbing, meaning good candidates exist somewhere in the sample set.
But Self-choice stays flat, so practical utility depends on verification and selection, not just generation diversity.
What is actually failing? Decision-loop decomposition
A general benchmark is most useful when it separates failure sources: intent inference, tool choice, interface syntax, memory over time, and answer verification. This section visualizes that decomposition.
Research takeaway One aggregate score hides whether an agent misunderstood the request, chose the wrong tool, or generated a correct answer and failed to recognize it.
Practical takeaway Product teams should evaluate routing, tool arbitration, state compression, and checking—not just end success.
Open question Some degradation may reflect true robustness limits; some may be wrapper unfairness. Better benchmarks disentangle both.
References
Compact source list used for this visual companion. arXiv IDs found in provided material: 2602.18998.
Li et al., 2026 Benchmark Test-Time Scaling of General LLM Agents arXiv:2602.18998