Sample one frozen LLM N times, score every sample by similarity to the others, output the top scorer. Step through it.
Click a node.
Click a milestone above.
Which ingredients does each method need?
Decision trees are swapped for LLM samples.
Each sample's score = sum of its similarity to every other sample. The similarity function is task-specific. Hover the cells.
Hover a cell or a row label.
Monte Carlo of plurality voting (mock answer distribution: one correct answer, five wrong ones). Raise error correlation to see the independence assumption fail.
Illustrative curves shaped to the reported endpoints (13B: 59% on GSM8K at N=40; 70B single-shot: 54%). Other tasks are mock data. Hover the chart.
Reported ranges: GSM8K +12–24, HumanEval +4–9, Chess and MATH single digits.
Mock values except the reported collapse: Debate + Agent Forest on HumanEval with both Llama2 models scores 0.
A synthetic summation task isolates three difficulty dimensions. Curves are schematic.
Properties of the three dimensions motivate voting at finer granularity.
Rough proxy: cost ∝ parameters × samples (70B ≈ 5.4 × a 13B call). Mock curves; ignores batching and memory effects.
Dots: gain at N=40 from the mock data. Dashed: illustrative extrapolation, never tested in the paper.
Discussion verdicts on what the evidence supports.