Interactive Visualization Princeton · 2026 arXiv:2604.11753 Original Viz Link

Agentic Aggregation for Long‑Horizon AI Tasks

Parallel rollouts are easy to launch. The hard part is finding the one buried clue inside long agent traces full of search queries, tool calls, observations, and half-finished plans. This page turns that idea into visual systems: trace heatmaps, adaptive inspection flows, and compute-vs-quality comparisons.

Benchmarks
6
Deep research, web search, navigation, software-style tasks
Model Families
3
Same family used for rollout and aggregator
Avg Gain
+5.3 pt
Absolute improvement over static aggregation baselines
Key Idea
Inspect traces
Search completed trajectories instead of voting on outputs
Core Bottleneck
Aggregation gets harder as trajectories get longer
Paper Abstraction
get_solution · search_trajectory · get_segment
Transcript arXiv IDs

References

Yao et al. 2023. Tree of Thoughts
Snell et al. 2024. Large Language Monkeys
Yao et al. 2023. ReAct
Snell et al. 2025. Best-of-N Test-Time Scaling
Shinn et al. 2023. Reflexion
Related benchmark context: BrowseComp, HLE, WebArena, SWE-bench