Real production websites stay intact. The agent navigates, types, uploads, and progresses through multi-step flows; only the final irreversible submission request is intercepted.
Each category mixes write-heavy form filling, uploads, auth, dynamic rendering, and anti-bot nuisance. Hover cells to inspect where failures likely cluster.
Toggle the baseline environment. The contrast is the point: models that look strong on stable benchmarks can underperform dramatically once interaction becomes closed-loop and messy.
Not all errors are “reasoning.” This stacked chart separates synthetic failure mass into perception/UI targeting, state tracking, timing, auth/access, and safety/website hostility.
Step through the interaction loop. Errors compound because the environment changes after every click, scroll, render, and modal.
Live tasks are hard to score from the final page alone. This matrix visualizes how multiple evidence sources jointly support a verdict.
Compact source map for the benchmark context and related web-agent evaluation work.