More Agents Is All You Need:
Sampling Beats Scaling

Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, Deheng Ye · Tencent · 2024 (TMLR) arXiv:2402.05120 Agent Forest

The Agent Forest pipeline

Sample one frozen LLM N times, score every sample by similarity to the others, output the top scorer. Step through it.

Lineage: from Random Forests to Agent Forest

Click a node.

Click a milestone above.

Agent Forest vs its precursors

Which ingredients does each method need?

Same intuition, new ensemble member

Decision trees are swapped for LLM samples.

Similarity-weighted voting

Each sample's score = sum of its similarity to every other sample. The similarity function is task-specific. Hover the cells.

Hover a cell or a row label.

Why voting works — and when it stops

Monte Carlo of plurality voting (mock answer distribution: one correct answer, five wrong ones). Raise error correlation to see the independence assumption fail.

Accuracy vs ensemble size

Illustrative curves shaped to the reported endpoints (13B: 59% on GSM8K at N=40; 70B single-shot: 54%). Other tasks are mock data. Hover the chart.

Gain at N=40 by task

Reported ranges: GSM8K +12–24, HumanEval +4–9, Chess and MATH single digits.

Stacking with other methods

Mock values except the reported collapse: Debate + Agent Forest on HumanEval with both Llama2 models scores 0.

What makes a task benefit? (Section 6)

A synthetic summation task isolates three difficulty dimensions. Curves are schematic.

Two derived variants

Properties of the three dimensions motivate voting at finer granularity.

The missing compute-matched baseline

Rough proxy: cost ∝ parameters × samples (70B ≈ 5.4 × a 13B call). Mock curves; ignores batching and memory effects.

Gains shrink as models get stronger

Dots: gain at N=40 from the mock data. Dashed: illustrative extrapolation, never tested in the paper.

Claim scorecard

Discussion verdicts on what the evidence supports.

References

  1. Li et al., More Agents Is All You Need, 2024.
  2. Wang et al., Self-Consistency Improves Chain of Thought Reasoning, ICLR 2023.
  3. Wei et al., Chain-of-Thought Prompting Elicits Reasoning, 2022.
  4. Wang et al., Rationale-Augmented Ensembles in Language Models, 2022.
  5. Du et al., Improving Factuality and Reasoning through Multiagent Debate, 2023.
  6. Kapoor et al., AI Agents That Matter, 2024.
  7. Lu et al., Blending Is All You Need, 2024.
  8. Wan et al., Knowledge Fusion of Large Language Models, 2024.