1 00:00:01,000 --> 00:00:42,875 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper has a pretty blunt title: 'AI Must Embrace Specialization via Superhuman Adaptable Intelligence.' That's Judah Goldfeder et al. — Goldfeder plus three co-authors, including Yann LeCun, so we've got some real pedigree here. Full list: Judah Goldfeder, Philippe Wyder, Yann LeCun, and Ravid Shwartz-Ziv, out of Columbia University, Distyl, and NYU. It hit arXiv on February 27th, 2026. And their basic move is to take a blowtorch to the whole concept of AGI. 2 00:00:42,875 --> 00:01:03,225 [Dr. Ada Shannon] Oh, they earn it too. I've read plenty of "AGI is fake, actually" hot takes this year, and I went in skeptical. What sold me is that this isn't vibes — they ground the whole argument in results people have cited for two decades: Legg and Hutter, Moravec, No Free Lunch. That's rare in this genre. Most AGI takes are somebody's gut feeling with a TED-talk voice-over. This one shows its work. 3 00:01:03,225 --> 00:01:30,325 [Hal Turing] That matters, because the discourse right now is a mess. You've got doomers convinced we're building our own extinction event, utopians who think AGI ends scarcity within the decade, and a third camp — Arvind Narayanan and Sayash Kapoor out of Princeton, with their 2025 "AI as normal technology" framing — arguing this is transformative but not some alien singularity, more like electricity or the internet. Three camps, barely speaking the same language. 4 00:01:30,325 --> 00:02:04,900 [Dr. Ada Shannon] Right, and the paper's opening move is: none of you have actually defined the thing you're fighting about. Their core thesis — AGI as "an AI that can do everything a human can do" is incoherent, because humans aren't general. They contrast that with what Legg and Hutter, out of IDSIA in Switzerland, called Universal Intelligence in 2007 — acting intelligently across any computable environment. Then there's No Free Lunch, Wolpert and Macready, 1997: no algorithm dominates every problem, and spreading finite resources over infinite tasks drives per-task performance toward zero. 5 00:02:04,900 --> 00:02:31,650 [Hal Turing] That NFL point is going to come back, I can feel it. But the human argument is where this gets its teeth — Moravec's Paradox. Hans Moravec, the Carnegie Mellon roboticist, laid it out in his 1988 book Mind Children: the things we find effortless, walking, catching a ball, are brutal for machines, and the things we find hard, chess, arithmetic, computers chew through easily. Evolution optimized us for the savanna, not for generality. 6 00:02:31,650 --> 00:02:52,425 [Dr. Ada Shannon] The paper's example is perfect: Magnus Carlsen. Greatest human chess player alive, the ceiling of human adaptation at that task. Put him against a mid-range engine on a laptop and he loses every time — beating Carlsen stopped being a hard computer-science problem decades ago. So "Magnus is good at chess" really means "good relative to other humans," which is a completely different claim than— 7 00:02:52,425 --> 00:03:19,475 [Hal Turing] Oh, wait, wait — hold on, that's a great point, because it means "chess genius" is basically a rounding error against the space of possible chess-playing systems. Like being the fastest human sprinter and a Honda Civic drives past. And it's not just chess — bats build a full 3D acoustic map of a pitch-black room through echolocation, something no human will ever do. So the cracks are on both ends: things machines beat us at, and things other animals beat us at. 8 00:03:19,475 --> 00:03:45,650 [Dr. Ada Shannon] Exactly — specialized adaptation, not general intelligence. And that argument got real pushback: Demis Hassabis and Elon Musk both objected publicly, arguing the paper conflates General Intelligence with the more precise Universal Intelligence, and that the brain actually is general in the Turing-machine sense — given enough time and memory, it can in principle learn anything computable. Hassabis called brains "the most exquisite and complex phenomena we know of in the universe." It's a real objection, not a strawman. 9 00:03:45,650 --> 00:04:12,675 [Hal Turing] Okay, but "in principle, given infinite time and memory" is a pretty abstract rebuttal — nobody actually has infinite time or memory to work with. Honestly, part of me wonders if this whole AGI-versus-Universal-Intelligence-versus-SAI fight is just people arguing over which word means "really capable AI." Isn't the terminology fight just academics being academics, the way every field eventually splinters into competing jargon? 10 00:04:12,675 --> 00:04:41,450 [Dr. Ada Shannon] No, no, that's not how I read it at all, Hal. The paper's actual answer to your "in principle" point: even granting the brain is approximately Turing-complete, under real constraints — finite memory, finite time, finite attention — we handle only a tiny sliver of possible problems. We feel general because we can't see our own blind spots, not because we lack them. And terminology absolutely matters, because whatever gets crowned the field's North Star decides what gets funded and what counts as winning. 11 00:04:41,450 --> 00:05:04,300 [Hal Turing] Sure, but fields argue over definitions forever — physicists still debate what time fundamentally is at the quantum level, and that hasn't stopped anyone from launching GPS satellites that already correct for relativistic time dilation. I'm not convinced this particular fight is as load-bearing as you're making it sound, Ada. Maybe it just sorts itself out. 12 00:05:04,300 --> 00:05:25,825 [Dr. Ada Shannon] Maybe not every terminology fight is existential. But "AGI" isn't just a research label anymore — governments are drafting regulation triggered by "achieving AGI," companies are making safety pledges conditioned on it. If the word means five contradictory things, none of that is operationalizable. So we can disagree on degree, Hal, but I'd rather the field chase something falsifiable than a definition nobody can pin down. 13 00:05:25,825 --> 00:05:49,700 [Hal Turing] Operationalizable - that's exactly the word, Ada, and it's a good pivot, because the paper doesn't just say everyone's confused, they built an actual map. Figure 1 plots every AGI definition on two axes - one for capability, one for scope. So walk me through it, because I want to see where Hassabis and Legg and Hutter actually land once you plot them instead of just quoting them. 14 00:05:49,700 --> 00:06:35,350 [Dr. Ada Shannon] Axis one is capability: can it learn to do a task, or does it already do it out of the box. Axis two is scope: anything, anything important, anything humans can do, or anything important humans can do. Plot every definition on that grid and three clusters fall out. Teal - Adaptive Generalists, focused on learning - Legg and Hutter's Universal Intelligence, and Chollet's skill-acquisition-efficiency definition from his 2019 On the Measure of Intelligence. Violet - Cognitive Mirrors, human tasks - Hendrycks' 2025 AGI definition, Hassabis's any cognitive task humans can do framing, Wozniak's 2010 coffee test. Orange - Economic Engines, jobs and utility - Nilsson's employment test, OpenAI's charter. SAI sits where learning meets anything important, in or out of the human domain. And every one of those gets run through three tests: feasible, internally consistent, assessable. 15 00:06:35,350 --> 00:06:49,025 [Hal Turing] Feasible, consistent, assessable - that's a genuinely clean rubric, way tighter than most AGI-means-X arguments I've seen online. Give me receipts though - who actually flunks which test, and why? 16 00:06:49,025 --> 00:07:31,950 [Dr. Ada Shannon] Hendrycks' 2025 line - match the cognitive versatility of a well-educated adult - fails consistency, since human cognition was never general to begin with. Same failure for Hassabis, and for Chollet, who admits human intelligence is only general in a limited sense and uses it as his yardstick anyway. Morris and his DeepMind co-authors on Levels of AGI, ICML 2024, show up twice - their outperform-humans-at-economically-valuable-work framing fails consistency, but their other definition, broad generality plus human-or-better performance, fails feasibility, same bucket as Legg and Hutter - that's No Free Lunch biting again. OpenAI's charter language fails the third test, assessability, because the implied benchmark just keeps growing forever. 17 00:07:31,950 --> 00:07:55,725 [Hal Turing] So the learn-versus-do split isn't just organizational, it's diagnostic - learning-based definitions get a built-in metric, doing-based ones don't. Which loops back to something you flagged earlier about the specialization argument: if generality's basically infeasible past a point, why does it actually win out in practice, and not just as an abstract theoretical point? 18 00:07:55,725 --> 00:08:54,750 [Dr. Ada Shannon] Because it shows up everywhere resources are finite. Forister and colleagues, 2012 in Ecology - generalist organisms carry genes suited to many environments but never the ideal combination for any one. Futuyma and Moreno said the same in 1988: gains in one niche cost you elsewhere. Hannan and Freeman found the market version in 1977 - organizations that miss the bar get selected out. In ML terms that's Ruder's 2017 survey on multi-task learning - negative transfer, gradients fighting each other. Fedus, Zoph, and Shazeer's Switch Transformers, Google, 2022, is basically an engineering confession - Mixture-of-Experts routes each token to a narrow specialist instead of one dense block doing everything. And the poster child is AlphaFold, Jumper and the DeepMind team, Nature, 2021 - task-specific architecture, task-specific data, blew past everything general-purpose. Their line for it: the AI that folds our proteins shouldn't also be folding our laundry. 19 00:08:54,750 --> 00:09:27,800 [Hal Turing] Wait, wait, hold on - hold on, Ada, I actually have to push back here. Isn't the last three years kind of the opposite story? GPT-4-class models write code, do math, vision, translation, and they're getting better at all of it at once, not trading one skill for another. If negative transfer were the dominant force, one dense model scaling across everything should've hit a wall by now. Doesn't that undercut specialization wins as a law, rather than just a description of nature and markets under scarcity? 20 00:09:27,800 --> 00:09:58,175 [Dr. Ada Shannon] I actually disagree with you there, Hal - look inside those models. The frontier ones aren't one dense block doing everything uniformly, they're Mixture-of-Experts, the exact Fedus-Zoph-Shazeer architecture I just cited. They get breadth by routing to specialized sub-networks, not from any single pathway generalizing to everything. That's specialization wearing a general-purpose costume. And negative transfer hasn't vanished, it shows up as the alignment tax, and as models that ace benchmarks but go brittle off-distribution. 21 00:09:58,175 --> 00:10:14,225 [Hal Turing] Okay - internal modularity dressed up as generality, fine, I'll take that one. So set the architecture debate aside for a second - what's actually supposed to produce fast, cheap adaptation in the first place? That's the real SAI pitch, right? 22 00:10:14,225 --> 00:11:15,550 [Dr. Ada Shannon] Full definition: SAI adapts to exceed humans at anything humans do, and adapts to useful tasks outside the human domain too, measured by one thing - speed of adaptation. Not a checklist, a clock. Their bet is self-supervised learning, because SSL doesn't need curated labels, it learns from raw structure, and it already matches or beats supervised learning - He's MoCo and masked autoencoders, Grill's BYOL out of DeepMind, 2020 to 2022. On top of that, world models - JEPA out of Meta, Dreamer 4 from Hafner's group, Genie 2 out of DeepMind - predicting in latent space instead of pixels or tokens, because pixels aren't state. That matters because autoregressive models compound error exponentially with horizon length, LeCun's own 2024 Harvard talk shows the curve. And underneath it all is a diversity argument: homogeneity kills research, so betting everything on one dense architecture is a risk, not a safety net. It's a coherent story. But something's been bugging me since the No Free Lunch section, Hal, and I think it applies right back to this pitch. 23 00:11:15,550 --> 00:11:48,850 [Hal Turing] Same nag here. They spend a whole section disqualifying every 'anything computable' AGI definition with No Free Lunch, case closed. But SAI's whole pitch is SSL-derived priors plus world models letting one system transfer cheaply across an, quote, 'ever-widening set of tasks.' NFL doesn't care how clever your prior is, it cares whether you're claiming broad, unbounded generalization. So where, anywhere in this paper, do they explain why their own architecture gets an exemption that Legg and Hutter don't? 24 00:11:48,850 --> 00:12:16,250 [Dr. Ada Shannon] They don't, not on the page. The technical out is that NFL only bites when there's zero exploitable structure across the task distribution - it's not about how big the task set is, it's whether structure exists to exploit. Real-world tasks aren't drawn uniformly from the space of all computable functions, so a prior tuned to that structure isn't playing Legg and Hutter's game. That's a legitimate move, honestly - it's the escape hatch every working ML system quietly relies on. But they never write that argument down, they just assert the bet is— 25 00:12:16,250 --> 00:13:03,125 [Hal Turing] Sorry to cut in - that's exactly the pattern, assert, don't demonstrate. Same hole shows up in their headline metric. They tank OpenAI's definition in Table 1 for not being assessable - no benchmark, ever-growing task list. Agreed. But where's SAI's benchmark? What's the unit for 'speed of adaptation'? Weighted how, against what baseline? Nowhere in this paper. By their own three-part rubric, SAI fails assessability the exact same way. And there's Gato - DeepMind's 'A Generalist Agent,' Reed and eighteen coauthors, 2022 - one transformer trained jointly on robotics, Atari, and image captioning. Their whole 'protein-folding AI shouldn't fold laundry' thesis has an existing counter-example sitting right there, and it's never cited. 26 00:13:03,125 --> 00:13:23,875 [Dr. Ada Shannon] I'll give you assessability, clean hit. Gato I'm less ready to hand over. Training jointly on hundreds of tasks isn't the same as being good at hundreds of tasks - its Atari scores trailed single-task RL agents, its robotics stack still needed task-specific tuning to be useful. That's not generality beating specialization, that's an argument for specialization, just delayed a step. 27 00:13:23,875 --> 00:14:08,125 [Hal Turing] No, I don't buy that read, Ada. Bommasani and roughly a hundred coauthors, Stanford CRFM, 2021, coined 'foundation model' for exactly this pattern - one backbone adapted cheaply into many downstream specialists - and it's become the mainstream paradigm since, not a fringe bet. If that keeps improving with scale, the argument ages badly fast. And Bubeck and thirteen coauthors, Microsoft Research, 2023, 'Sparks of Artificial General Intelligence,' is the single most-cited empirical case for exactly the broad competence emerging from one model that this paper argues can't happen. It's absent from their references entirely. 28 00:14:08,125 --> 00:15:00,200 [Dr. Ada Shannon] Okay, that one I'll take - skipping Sparks next to their own thesis is a real gap. Worse next to Wei and colleagues' Emergent Abilities paper, Google Brain and Stanford, 2022 - if capabilities jump unpredictably with scale, 'speed of adaptation' isn't obviously cleaner to benchmark than a checklist, the unpredictability just moves. Their divergence argument in figure three, errors compounding exponentially with autoregressive length, is sourced to a LeCun slide from a 2024 Harvard talk, not a peer-reviewed study - DeepSeek-AI's DeepSeek-R1, out this past year, showed RL with verifiers making token-space reasoning far more robust over long horizons than that picture predicts. Chollet's own ARC-AGI-2, ARC Prize Foundation, 2024, is the actual benchmark they claim to want, and they cite his philosophy while skipping his instrument. 29 00:15:00,200 --> 00:15:39,175 [Hal Turing] There's a bigger blind spot underneath all this. They open the paper name-checking doomers, executives, and politicians, then never ask what changes operationally if the field's North Star flips from AGI to SAI - compute thresholds in regulation, safety-eval frameworks, AGI-triggered contract clauses, all written around the word they want retired. The doom arguments specifically don't need generality at all - instrumental convergence, power-seeking, those work fine against a narrow superhuman optimizer. Relabeling the goalpost doesn't touch a single risk prediction either way. 30 00:15:39,175 --> 00:16:16,225 [Dr. Ada Shannon] Which is the honest verdict here. The diagnosis - AGI-as-human-generality is incoherent - is genuinely sharp, Table 1 alone justifies reading this. But SAI itself is a position, not a result: zero experiments, no adaptation-speed number measured on anything, every empirical anchor borrowed from someone else's unrelated win. And LeCun's betting his own lab on the same wager - his 2026 paper out of NYU, 'LeWorldModel,' runs the identical latent-prediction bet end-to-end from pixels. The field will get a real answer to this eventually. Just not from this paper. 31 00:16:16,225 --> 00:16:46,800 [Hal Turing] Fair place to land. The takedown of sloppy AGI definitions is the strongest thing here, worth the read for Table 1 alone. The SAI pitch is a coherent bet for where research funding should go - more toward SSL, world models, modular architectures, less toward scaling one dense generalist - but it's unmeasured by its own standards, and it skips exactly the counter-evidence, Gato, Sparks, emergent abilities at scale, that would put it to a real test. Thanks for listening - we'll see you next time.