ClawBench for Real-World Online AI Agents

A visual companion for the episode: how benchmark scores collapse when browser agents leave deterministic sandboxes and face the live web’s cookie banners, dynamic layouts, auth friction, and write-heavy workflows.
arXiv:2604.08523 153 tasks 144 live websites 15 categories 7 frontier models Real-browser evaluation
Core tension
Realism ↔ Safety ↔ Reproducibility
Lead score
33.3%
Best prior-style scores
67–75%
Big story
Strong LLM ≠ strong agent
Mocked visual framing uses paper-reported headline numbers plus synthetic category/friction detail for exploration.

1. Live-web benchmark pipeline

Real production websites stay intact. The agent navigates, types, uploads, and progresses through multi-step flows; only the final irreversible submission request is intercepted.

agent/browser interaction safety interception evaluation evidence real-web friction points

2. Task landscape and friction profile

Each category mixes write-heavy form filling, uploads, auth, dynamic rendering, and anti-bot nuisance. Hover cells to inspect where failures likely cluster.

Heatmap values are synthetic but grounded in the episode’s qualitative claims: travel, job applications, finance, and admin flows are more brittle because a single bad field or auth state can invalidate the whole task.

3. Sandbox confidence vs live-web collapse

Toggle the baseline environment. The contrast is the point: models that look strong on stable benchmarks can underperform dramatically once interaction becomes closed-loop and messy.

4. Failure budget decomposition

Not all errors are “reasoning.” This stacked chart separates synthetic failure mass into perception/UI targeting, state tracking, timing, auth/access, and safety/website hostility.

5. Closed-loop agent difficulty map

Step through the interaction loop. Errors compound because the environment changes after every click, scroll, render, and modal.

state observation policy/memory action execution common breakpoints

6. Evaluator evidence matrix

Live tasks are hard to score from the final page alone. This matrix visualizes how multiple evidence sources jointly support a verdict.

References

Compact source map for the benchmark context and related web-agent evaluation work.