AI Post Transformers — Episode Companion

Superhuman Adaptable Intelligence Challenges the Idea of AGI

Humans aren't generally intelligent, so "an AI that can do everything a human can do" isn't a coherent target. A two-axis map of every major AGI definition, a three-part rubric that fails most of them, and the counter-evidence the paper's own Specialization-via-Adaptable-Intelligence pitch skips citing.

arXiv:2602.23643 Goldfeder, Wyder, LeCun, Shwartz-Ziv · 2026 Columbia · Distyl · NYU

Figure 1, reconstructed: where every AGI definition actually lands

Plot every definition on two axes — capability (can it learn a task, or must it already do it out of the box) and scope (anything, anything important, anything humans can do, or anything important humans can do) — and three clusters fall out. Hover a point for the definition and its source.

Mock axis placement reconstructed from the episode discussion of Figure 1 — not a pixel-accurate reproduction of the paper's figure.

Table 1, reconstructed: feasible / consistent / assessable

Every definition gets run through three tests. Green passes, orange fails. The rightmost column is an illustrative "argument centrality" heat value — how much weight the episode gave that failure — built with the same heatColor() gradient used elsewhere on this page.

Magnus Carlsen vs. a laptop

"Good at chess" collapses into "good relative to other humans" the instant you widen the comparison set. Illustrative Elo, not tournament data.

Human vs. bat spatial sense

The cracks run both directions — machines beat us at chess, bats beat us at building a 3D map of a pitch-black room. Illustrative accuracy, not a measured study.

Dense generalist vs. Mixture-of-Experts

Frontier models look general from the outside, but Fedus, Zoph & Shazeer's Switch Transformers architecture gets breadth by routing tokens to narrow specialists, not from one pathway generalizing to everything. Toggle the routing mode.

The SAI pipeline: self-supervised priors → world models → fast adaptation

SSL doesn't need curated labels; world models predict in latent space instead of pixels or tokens, because pixels aren't state.

Why latent prediction, not autoregression

Autoregressive models compound error exponentially with horizon length (LeCun's 2024 Harvard-talk chart). Toggle each line — illustrative curves, not measured data.

The unstated No-Free-Lunch exemption

NFL only bites when a task distribution has zero exploitable structure. The paper never writes down why SAI's priors get an exemption that Legg & Hutter don't — toggle to see the escape-hatch argument the hosts had to supply themselves.

The counter-evidence SAI's own rubric would flunk it on

By the paper's own three-part test, SAI's "speed of adaptation" has no benchmark, no unit, no baseline. Click a bar for the specific gap and the counter-citation missing from the paper's reference list.

References