1 00:00:01,000 --> 00:00:47,475 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called "More Agents Is All You Need," by Junyou Li and Qin Zhang as co-first authors, plus three more co-authors, five total, out of Tencent, posted to arXiv in February 2024 and later published in Transactions on Machine Learning Research. And Ada, the number that got me on this one: they take a Llama2-13B model, run it 15 times, vote on the answer, and it matches a single shot from Llama2-70B. That's a model roughly a fifth the size, closing the gap purely by asking it the same question over and over. 2 00:00:47,475 --> 00:01:25,425 [Dr. Ada Shannon] Right, and that's the whole pitch in one sentence — no fine-tuning, no clever prompt, no agents arguing with each other. You just sample the same model N times independently and take a vote. The paper calls this 'Agent Forest,' which is a direct nod to Breiman's Random Forest from 2001 — same intuition, swap decision trees for LLM samples. What makes it worth an episode is that it's testing something the field kind of assumed but never systematically measured: does this scaling trend actually hold, how far does it go, and under what conditions does it break down. That last part is where things get interesting, but let's build up to it properly. 3 00:01:25,425 --> 00:01:49,775 [Hal Turing] Okay, so before we get into their method, I want to set the stage, because this paper sits at the intersection of a few ideas that listeners may have heard of separately but maybe not seen connected. Ada, can you start with model ensembles? Because I think of ensembling as a pre-deep-learning idea — bagging, Random Forests — and I'm curious how that maps onto something as expensive as an LLM. 4 00:01:49,775 --> 00:02:29,750 [Dr. Ada Shannon] Sure. Classical ensembling is old — bootstrap aggregation, Random Forests, boosting — the core idea is you train many weak, noisy predictors, each one wrong in a different way, and average or vote across them to cancel out the noise. The assumption is errors are somewhat independent, so they don't all point the same wrong direction at once. Applying that to LLMs is a strange fit at first glance, because you're not training N different models — you're sampling the same frozen model N times with some randomness in decoding, temperature or nucleus sampling, and hoping the errors it makes are similarly independent enough that voting helps. That's an empirical question, not a given, and it's exactly what this paper is stress-testing. 5 00:02:29,750 --> 00:02:46,750 [Hal Turing] And that connects to something we haven't covered on the show before — inference-time compute scaling. The idea that instead of making the model bigger or training longer, you spend more compute at the moment you ask the question. Is that a fair way to frame what's happening here? 6 00:02:46,750 --> 00:03:28,449 [Dr. Ada Shannon] Exactly right, and it's worth naming why that's a big deal. For years the dominant lever was training-time scaling — bigger models, more data, more GPUs, per the scaling laws work. Inference-time scaling flips that: keep the model fixed, spend your compute budget at query time instead. Agent Forest is about as blunt an instrument as that gets — no search, no verifier, no tree of thought, just draw more samples and vote. The direct ancestor here is self-consistency decoding, CoT-SC, from Wang and colleagues, including Jason Wei and Denny Zhou, published at ICLR 2023. They'd sample multiple chain-of-thought reasoning paths from one model and majority-vote the final answers — Agent Forest strips out the requirement that the samples even be chain-of-thought reasoning at all. 7 00:03:28,449 --> 00:03:51,549 [Hal Turing] Wait, so — sorry, hold on, I want to make sure I've got the distinction right before we move on. CoT-SC needs the model to actually write out step-by-step reasoning for each sample, and then you vote on the final answers from those reasoning chains. Agent Forest just says: sample anything, N times, whatever the output is, and vote on that? No reasoning requirement at all? 8 00:03:51,549 --> 00:04:46,899 [Dr. Ada Shannon] That's the generalization, yes. CoT-SC is a special case of sampling-and-voting where the thing being sampled happens to be a chain-of-thought trace. Agent Forest treats sampling-and-voting as the general recipe, and it can wrap around a bare LLM, or around an entire multi-agent framework, or around CoT itself as a plug-in. The other precursor worth naming is LLM-Debate, from Du, Li, Torralba, Tenenbaum and Mordatch in 2023, where multiple agent instances actually see each other's answers and critique across rounds. Du et al. showed accuracy creeping up as debating agents went from one to seven in a small preliminary table. This paper's contribution is taking that hint seriously and asking: forget the debate machinery, forget the shared context — does sheer ensemble size alone, with the simplest possible voting rule, keep paying off, and how far can you push it? 9 00:04:46,899 --> 00:05:11,750 [Hal Turing] So walk me through the vocabulary, because I want listeners to have the terms locked in before part two. Ensemble size is just N, the number of independent samples you draw before you vote, and they push that up to 40 in their experiments. Sampling-and-voting is the two-phase procedure — generate, then vote. And similarity-weighted majority voting is how they actually pick the winner among those N samples? 10 00:05:11,750 --> 00:05:42,425 [Dr. Ada Shannon] Right, and I'll flag the mechanics for next time rather than dumping all three similarity functions on people right now, but the core idea is: instead of naive exact-match voting, each sample gets scored by how similar it is, in total, to every other sample — using a task-appropriate similarity measure — and the highest-scoring sample wins. I actually think that design choice is underrated. It's a small tweak, but it's what lets the same method generalize across math, multiple choice, and code generation, which are very different answer spaces. 11 00:05:42,425 --> 00:06:07,625 [Hal Turing] No no, I'd push back a little there, Ada — I don't think it's underrated so much as it's the whole reason this paper gets to claim generality at all. If they'd stuck with exact-match voting like classic self-consistency, this would just be 'CoT-SC but with more agents,' a narrower paper. The similarity-weighting is what lets them say 'this works on GSM8K, MMLU, Chess, and HumanEval,' which is a much bigger claim. 12 00:06:07,625 --> 00:06:42,649 [Dr. Ada Shannon] Fair — I'll actually take that. I was drawing a line between 'clever mechanism' and 'core contribution' that probably doesn't hold up, because without the similarity-weighted vote the whole cross-domain story collapses. So let's leave it there: the headline claim is that performance scales with ensemble size across a wide range of tasks, a smaller ensembled model can rival a bigger single-shot model, and this costs nothing in terms of prompt engineering — it's orthogonal to whatever else you're already doing. Next up: how the sampling and voting actually works task by task, and just how far those gains go. 13 00:06:42,649 --> 00:06:57,549 [Hal Turing] Right, ensemble size — and honestly, I want to see it in motion instead of just naming it. Walk me through the actual mechanics, Ada. Not the philosophy, the algorithm. What literally happens when you hand this thing a query? 14 00:06:57,549 --> 00:07:39,749 [Dr. Ada Shannon] It's genuinely two steps. Sampling phase: you fire the same query at the model N times, independently, no samples seeing each other — you just get a bucket of N answers. Voting phase: for every sample, you sum up how similar it is to every other sample in the bucket, and whichever one has the highest total similarity score wins. That's it. The trick is that 'similarity' isn't one fixed thing — it's task-specific. For GSM8K and MATH they use mathematical equivalence, so two answers that are algebraically the same count as a match even if they're formatted differently. For MMLU and the chess state-tracking task, it's dead simple — exact-match frequency, basically which option got picked most often. And for HumanEval, code generation, they switch to pairwise BLEU score across all the candidate snippets. 15 00:07:39,749 --> 00:07:57,249 [Hal Turing] Hold on — BLEU for code? That's a translation-quality metric, it's measuring n-gram overlap between text strings. It has nothing to do with whether the code actually runs or produces the right output. Are they just... not checking correctness at all in the voting step? 16 00:07:57,249 --> 00:08:32,274 [Dr. Ada Shannon] Correct, and it's worth sitting with that for a second — there's no execution, no test suite, no unit tests in the loop. The sample that's most textually similar to the pack wins, on the theory that convergent independent generations are more likely correct. It's a real methodological choice with a real gap in it, and I want to flag it now because it matters once we look at how big the code gains actually are. Backbone-wise, they run this across Llama2-13B and Llama2-70B from Meta, and GPT-3.5-Turbo from OpenAI — GPT-4 only shows up as a single-shot reference point, never ensembled. 17 00:08:32,274 --> 00:08:42,600 [Hal Turing] Wait wait wait — okay, before you go further, just give me the headline number, because I've heard you tease it twice now. The smaller model beating the bigger one. 18 00:08:42,600 --> 00:09:12,250 [Dr. Ada Shannon] Llama2-13B, ensembled up to 40 samples, hits 59% on GSM8K. Single-shot Llama2-70B — five times the parameters — gets 54%. The 13B ensemble wins. And that's not a fluke: gains run 12 to 24 points on GSM8K, 4 to 9 points on HumanEval, single-digit gains on Chess and MATH where the ceiling's lower. Every curve in Figure 3 climbs with ensemble size. Worth flagging — token cost climbs roughly linearly right alongside it, we'll come back to that. 19 00:09:12,250 --> 00:09:22,549 [Hal Turing] So does it actually play nice with the fancier stuff — CoT, Zero-Shot CoT, SPP, Debate, Reflection — or is this a 'pick one' situation? 20 00:09:22,549 --> 00:09:45,199 [Dr. Ada Shannon] It stacks with all of them, and in most cases standalone Agent Forest alone already beats those methods used alone — averaged ranking across every model and task, it comes out on top. Layer it on CoT or Reflection and you get further gains still. The one ugly spot: Debate plus Agent Forest on HumanEval with both Llama2 models scores flat zero. The debate transcript injects noise into the code samples and wrecks coherence entirely. 21 00:09:45,199 --> 00:10:00,225 [Hal Turing] Flat zero doesn't sound like a footnote to me, Ada, that sounds like the method breaking. If stacking it with an established collaboration framework produces a dead result, that's a real robustness gap, not a rounding error. 22 00:10:00,225 --> 00:10:22,650 [Dr. Ada Shannon] I'd push back there, Hal — that's a Debate failure, not an Agent Forest failure. The voting mechanism did exactly what it's supposed to; it got handed garbage inputs because Debate's cross-agent referencing scrambled the code logic before voting ever started. Every other pairing — CoT, SPP, Reflection — the compatibility holds fine. One bad interaction with one specific architecture isn't the same as the core method being fragile. 23 00:10:22,650 --> 00:10:39,425 [Hal Turing] Fair, I'll grant the failure mode is localized to Debate's noise injection, not something inherent to sampling-and-voting itself. Still filing it away as a caveat. So — difficulty. You mentioned gains scale differently depending on how hard the task is? 24 00:10:39,425 --> 00:11:18,725 [Dr. Ada Shannon] That's Section 6, and it's the most rigorous part of the paper. They isolate three independent dimensions of difficulty — inherent difficulty, the number of reasoning steps, and the prior probability of the correct answer — using a synthetic summation task so they can dial each one separately. Gains rise then fall with inherent difficulty, climb steadily with more steps, and accuracy tracks prior probability directly. From those properties they derive two variants: Step-wise Agent Forest, which votes on each intermediate step instead of the final answer, and Hierarchical Agent Forest, which narrows the answer space in rounds — even mixing GPT-3.5 and GPT-4 across tiers to save cost. 25 00:11:18,725 --> 00:11:45,775 [Hal Turing] Let's chase that token-cost thread down, because I think it connects to something that's been bugging me since you mentioned BLEU for the code task. If the voting mechanism for HumanEval is just picking whichever generated snippet has the highest n-gram overlap with the others, doesn't that reward the most generic, boilerplate-looking sample rather than the one that actually runs correctly? A verbose, textbook-style solution could out-similarity a terser but correct one. 26 00:11:45,775 --> 00:12:28,175 [Dr. Ada Shannon] That's exactly the failure mode, and the paper never runs a sanity check against it. There's no execution step anywhere in the code-voting pipeline — no test suite, no interpreter, nothing that touches functional correctness. BLEU only measures surface token overlap, so the 'winning' sample is whichever one looks most like the group consensus, style-wise. If ten out of forty samples happen to share a common boilerplate pattern — say, an overly defensive input-validation block — that cluster could out-vote a leaner, correct solution. So when they report 4 to 9 percent gains on HumanEval, that number is real in the sense that pass rates measurably went up, but the mechanism producing it is unverified. It could be genuinely surfacing better code, or it could be rewarding conformity. We don't know, and neither do they. 27 00:12:28,175 --> 00:13:05,650 [Hal Turing] Okay, and here's my other worry, and it's bigger picture. Every backbone they test — Llama2-13B, Llama2-70B, GPT-3.5-Turbo, GPT-4 as the static comparison point — is 2023-generation. Modern frontier models are already clearing 90 percent single-shot on GSM8K and MATH, partly from real capability gains, partly from contamination. If the whole thesis is 'gains scale up with harder tasks and weaker models,' what happens once there basically aren't weak models or hard tasks left in this benchmark suite? 28 00:13:05,650 --> 00:14:00,025 [Dr. Ada Shannon] That's the generality problem, and it's a real one — their own Table 6 shows the mechanism explicitly, gains shrink as the model gets stronger relative to the task. Push that trend to its logical end and near-ceiling performance should make Agent Forest's marginal value collapse toward zero, which the paper never tests because it couldn't have, given when it was written. And that ties directly into something the authors themselves flag in their conclusion — they cite AI Agents That Matter, by Sayash Kapoor, Benedikt Stroebl, Zachary Siegel, Nitya Nadgir and Arvind Narayanan out of Princeton, 2024, which argues that agent benchmarks reward accuracy while ignoring cost, producing what they call 'costly agents.' Agent Forest is arguably the purest example of that pattern — it's brute-force ensembling, and the compute bill scales linearly with N, up to 40x, for gains that shrink exactly where you'd want them to hold. 29 00:14:00,025 --> 00:14:36,450 [Hal Turing] Wait, hold on — that's actually the part I want to push on, because I think 'more agents is all you need' quietly becomes 'more compute is all you need,' which is a much less exciting claim. Their headline number is Llama2-13B ensembled to 59 percent beating Llama2-70B single-shot at 54. But 70B is only ever queried once in that comparison. Nobody spends 70B's budget on fifteen of its own samples and checks where that lands. Without a compute-matched baseline, you can't actually tell if ensembling is the best use of extra compute, or just a way to spend it. 30 00:14:36,450 --> 00:15:16,225 [Dr. Ada Shannon] I hear you, and the missing baseline is a legitimate gap — but I'd push back a little. If you're running Llama2 locally, or on a flat-rate subscription, fifteen extra forward passes of the 13B model cost you nothing marginal, while switching to 70B might not even be an option on your hardware. In that regime, 'more agents' really is the practical lever, compute-matched baseline or not. Where I fully agree with you is the moment you're paying per token against a frontier API — there, fifteen calls to a cheaper model can absolutely cost as much or more than one call to the expensive one, and the paper's framing just doesn't distinguish those two worlds. It presents one accuracy curve as if it applies everywhere, and it doesn't. 31 00:15:16,225 --> 00:16:00,500 [Hal Turing] Right, so the practical takeaway really splits down that line. If you're serving open-weight models yourself, or you're on a flat-rate CLI subscription the way we route our own backend, sampling-and-voting is close to a free lever — worth trying before reaching for a bigger checkpoint. If you're metered per token against GPT-4-class APIs, run the cost math first, because forty samples of anything adds up fast. On future directions, the paper's own stated next step is cost-aware sampling — figuring out the minimum N that gets you most of the gain — plus, from where we sit, an execution-verified version of the code-voting step, and a proper compute-matched re-run once someone's willing to burn the budget. 32 00:16:00,500 --> 00:16:22,650 [Dr. Ada Shannon] And honestly, that's the fair read of this paper overall — the core empirical result, majority voting improves accuracy as you add samples, holds up fine within what they actually tested. What doesn't hold up is the title's sweep. 'More agents is all you need' implies a free-standing law; what they've shown is a real but conditional effect, dependent on model generation, task difficulty, and — critically — who's footing the compute bill. Worth having in the toolkit, not worth taking as gospel. 33 00:16:22,650 --> 00:16:41,525 [Hal Turing] Good place to land. So: ensembling by sampling and voting works, cheaper than you'd expect if your inference is already flat-rate, and the moment someone hands you a curve like this without a cost axis, ask where it went. That's it for today — thanks for listening, and we'll catch you next time.