TUMIX Multi-Agent
Test-Time Scaling with Tools

Visual companion for the episode: a single strong LLM is turned into a specialist team with distinct tool-use policies — text, code, search, hybrids — then iteratively refined under roughly cost-matched inference budgets.

arXiv 2510.01279 Paper year 2025 Focus Inference-time compute Benchmarks HLE · GPQA · AIME Interactive viz link
Episode lens

The central comparison is not just more samples vs fewer samples, but homogeneous reasoning paths vs diverse computational pathways. The visuals below emphasize tool-policy diversity, iterative cross-agent revision, cost/latency tradeoffs, and where the gains likely come from.

Text-only reasoning Code execution Search / retrieval Hybrid agents Adaptive stopping Auto-designed agent rosters

Headline mock metrics from the paper discussion

+10.2
HLE gain on Flash
~49%
cost after early stop
+1.2%
auto-design bump
tool diversity
iteration / refinement
cost / latency pressure

From one LLM to a specialist panel

Switch between a single-agent baseline and the TUMIX multi-agent orchestration. Hover nodes and links for local details.

Tool-policy diversity matrix

Agents differ less by weights than by permissions, prompting, and willingness to call tools. Explore task classes and the expected utility of each pathway.

Performance under roughly cost-matched settings

Bars show benchmark performance from the episode’s cited figures. Toggle model family and compare accuracy, token cost, and latency pressure.

Where TUMIX sits in the inference-time landscape

This map situates TUMIX among self-consistency, tool-using agents, executable reasoning, and agent-orchestration systems.

Compact references

Primary citations and neighboring work mentioned in the episode.

TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture
Chen et al., 2025
arXiv:2510.01279
PAL: Program-aided Language Models
Gao et al., 2022
scholar link
ReAct: Synergizing Reasoning and Acting
Yao et al., 2023
scholar link
Mixture-of-Agents Enhances LLM Capabilities
Wang et al., 2024
scholar link
Scaling LLM Test-Time Compute Optimally...
Brown et al., 2024
scholar link
Program-of-Thoughts Prompting
Chen et al., 2022
scholar link
Self-MoA / Symbolic-MoE / DEI / SciMaster / GSA
2025 comparison baselines
baseline cluster
Why Do Multi-Agent LLM Systems Fail?
2024–2025 failure analysis
scholar link