1 00:00:01,000 --> 00:00:43,274 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called Why AI Systems Don't Learn and What to Do About It: Lessons on Autonomous Learning from Cognitive Science. It's by Emmanuel Dupoux, Yann LeCun, and Jitendra Malik — three authors, spanning FAIR at Meta, the École des Hautes Études en Sciences Sociales, NYU, and UC Berkeley. It went up on arXiv in mid-March 2026. And Ada, the line that stopped me cold is right in the introduction: once you deploy an AI model, it learns essentially nothing. Zero. If the world shifts under it, you don't retrain it — you rebuild it, by hand. 2 00:00:43,274 --> 00:01:14,049 [Dr. Ada Shannon] Right, and that's the consequence you should sit with before we even get to their proposed fix: every model you've ever used in production is, cognitively speaking, dead on arrival. It does inference forever, but the learning stopped the moment the training run ended. Compare that to literally any toddler, who is running a live experiment on the world every waking second. The paper's framing is that this isn't a minor gap — it's the roadblock standing between where deep learning is now and anything we'd honestly call autonomous intelligence. 3 00:01:14,049 --> 00:01:40,750 [Hal Turing] That toddler comparison is doing a lot of work in the paper, and I love it. A kid with a new toy will bang it around to see what happens, that's learning through action. Then they'll watch another kid play with it and copy the gesture, that's learning through observation. They'll ask a caregiver how it works, or just sit and stare into space imagining what else it could do. Four different learning modes, and the kid switches between them fluidly, in real time, with nobody scheduling it. 4 00:01:40,750 --> 00:02:16,724 [Dr. Ada Shannon] Whereas in AI, those modes are siloed into entirely separate subfields — self-supervised learning, supervised learning, reinforcement learning — each with its own data pipeline and its own team of humans deciding when to use it. The paper calls this the externalization of learning: the switching logic that should live inside the system instead lives inside an MLOps pipeline run by people. And this connects to a critique that's been circulating loudly this year — Silver and Sutton's Era of Experience paper — that today's models are starved of real interaction and are about to hit a data wall of usable text. 5 00:02:16,724 --> 00:02:36,349 [Hal Turing] Okay, before we get to their fix, walk me through the vocabulary, because the paper builds two big conceptual buckets before System M ever shows up, and I think listeners need these locked in first. What's the actual fault line they're drawing between how a model watches the world versus how it acts on it? 6 00:02:36,349 --> 00:03:09,975 [Dr. Ada Shannon] Right, they call it System A and System B. System A is observation-based learning — the organism, or the model, just passively sits there building a statistical model of whatever sensory input flows past it. That's the umbrella for basically all self-supervised learning, across text, images, audio, whatever modality — a masked-language model or a contrastive image encoder is a System A learner. It's powerful, it scales beautifully with data, but on its own it has no way to decide what to look at next, and its representations are never grounded in anything the system actually does. 7 00:03:09,975 --> 00:03:12,750 [Hal Turing] And I'm guessing System B is— 8 00:03:12,750 --> 00:03:44,550 [Dr. Ada Shannon] —the flip side, exactly. System B is action-based learning: an agent that acts on the world, gets feedback, and adjusts. That's the paper's umbrella for reinforcement learning, control, planning — anything where you're optimizing behavior against reward rather than just modeling input statistics. It's grounded, it can find genuinely novel solutions through search, but it's brutally sample-inefficient and falls apart once the action space gets large and the reward signal gets murky, which is most of real life. 9 00:03:44,550 --> 00:04:02,400 [Hal Turing] So neither one alone gets you a toddler — System A can't tell you what to look at next, and System B is too sample-hungry to figure out where to look on its own. Which is presumably where their third piece comes in, this System M thing you keep hinting at. What is it actually doing? 10 00:04:02,400 --> 00:05:08,400 [Dr. Ada Shannon] That's their proposed orchestrator, and the analogy they reach for is genuinely clever: a software-defined-networking control plane, sitting alongside System A, System B, and an episodic memory buffer. Picture the plain arrows in their diagram as the fat pipes — raw sensory streams, motor commands, latent codes — and System M never touches any of that directly. It only watches a handful of low-bandwidth channels they call meta-states: epistemic signals like prediction error or confidence, species-specific signals like a pre-wired startle response to a looming shape, and, for embodied agents, somatic signals like energy levels or pain. Based on that telemetry it issues meta-actions — literally opening or closing which pipes are connected — so a training pipeline assembles itself on the fly instead of a human engineer wiring one up in advance. It's explicitly meant to automate the job a human MLOps engineer does today by hand, and they lean directly on the Kreutz et al. software-defined-networking survey from 2015, just applied to cognition instead of packets. They trace this whole architecture back to LeCun's own 2022 proposal for autonomous machine intelligence — this paper is very much the sequel to that one, now with Dupoux and Malik pulling in the cognitive-science side. 11 00:05:08,400 --> 00:05:28,950 [Hal Turing] Okay, so if M is the thing deciding when System B gets to go explore versus when System A gets to sit and consolidate, that routing itself is a policy. Does M learn that policy the same way System B learns one — gradient descent, reinforcement, something adaptive — over the agent's actual lifetime? 12 00:05:28,950 --> 00:06:02,350 [Dr. Ada Shannon] That's the one place they explicitly break their own pattern, and it surprised me too. They propose that System M's core routing policy is hardwired — their phrase is 'an evolutionarily fixed transition table' — dictating when to explore, when to plan, when to act. Not learned during the organism's life at all. Every other component in this paper is about learning, and here they're saying the switch that governs learning is itself not learned, it's baked in before the organism is ever born. Strong claim, and it's going to matter a lot once we get to how they propose actually building this thing, because it shifts the whole problem from 'train M' to 'evolve M.' 13 00:06:02,350 --> 00:06:12,500 [Hal Turing] Let's get concrete on the machinery for a second before we go there — strip System A down to the actual mechanics. What's literally happening inside it? 14 00:06:12,500 --> 00:06:53,150 [Dr. Ada Shannon] It's almost disappointingly simple on paper. You've got raw samples coming from some distribution, a task generator they call G that takes a sample and carves it into an input and a target, and then you train a representation to minimize a loss between its prediction and that target. That's the entire self-supervised learning family in one abstraction — masked language modeling, contrastive image pretraining, all of it fits this shape. But look at what it needs: a hand-built dataset and a hand-built G, both requiring real domain expertise. There's no built-in mechanism for deciding what data to go get next, the representations are disconnected from the agent's ability to act, and because it's purely observational, it can't tell correlation from causation. 15 00:06:53,150 --> 00:06:56,000 [Hal Turing] And System B — same treatment? 16 00:06:56,000 --> 00:07:34,250 [Dr. Ada Shannon] System B is the classic control-and-RL formulation: an agent in some state, taking an action, the world evolves under a transition function, the agent collects a reward, and it's trying to maximize expected discounted return over a horizon. If you already know the world's dynamics, you don't even need to learn anything — that's control theory. If you don't, that's reinforcement learning and planning. The limitation is the flip side of System A's: it's notoriously sample-inefficient, it falls apart in high-dimensional or open-ended action spaces, and it needs a well-specified reward function and interpretable actions, which basically don't exist outside toy domains. 17 00:07:34,250 --> 00:07:40,625 [Hal Turing] So where do they actually show these two propping each other up, with real systems people have built? 18 00:07:40,625 --> 00:08:30,125 [Dr. Ada Shannon] This is where the paper gets genuinely useful. System A helps System B by compressing raw pixels into tractable state and action representations, and by learning predictive world models — MuZero, Schrittwieser et al. out of DeepMind, 2020, and Dreamer, Hafner and colleagues largely at Google Brain, 2020, both turn blind trial-and-error into actual planning. V-JEPA, Bardes et al. out of Meta FAIR, 2024, does the same in latent space for physical intuition. System A can also hand System B intrinsic reward — prediction error as a built-in curiosity signal. Going the other way, System B helps System A through active SSL, where the agent's own gaze picks the most informative slice of data to learn from, and goal-directed SSL, where just pursuing a task throws off grounded data as a byproduct. It's Gibson's 1966 line — we see in order to move, and we move in order to see. 19 00:08:30,125 --> 00:08:47,475 [Hal Turing] But that's a chicken-and-egg problem, isn't it? System A needs action-grounded data from System B, System B needs perceptual structure from System A, and M needs calibrated signals from both of them to know what to route. How does any of this ever get off the ground? 20 00:08:47,475 --> 00:09:20,275 [Dr. Ada Shannon] Right, and this is exactly the problem Section 4 is built to answer, by borrowing the evolution-versus-development split from biology. There's an inner loop, the developmental scale, where System A and System B update their own parameters through interaction with the environment, with M held completely fixed. Then there's an outer loop, the evolutionary scale, that optimizes a set of meta-parameters — the initial architecture, M's transition table — against a fitness function measured across the agent's entire life cycle. 21 00:09:20,275 --> 00:09:29,200 [Hal Turing] Oh wait wait wait — so one entire simulated life, start to finish, is just one data point for that outer optimization? 22 00:09:29,200 --> 00:10:12,350 [Dr. Ada Shannon] Exactly, and they don't dodge how brutal that is. They state plainly that optimizing the outer loop requires running millions of simulated life cycles, each of which itself involves learning over millions of datapoints, and that extending bilevel optimization to architectures this large has severe, well-documented scalability issues. They're not claiming this is solved. What they do point to is existing work already chipping at pieces of it in constrained settings — MuZero and Dreamer for the world-model half, Video Pretraining, Baker et al. out of OpenAI, 2022, for learning latent actions straight from video, and Prioritized Experience Replay, Schaul et al. at DeepMind, 2016, alongside standard active learning for the input-selection half. None of those pieces add up to a System M yet. 23 00:10:12,350 --> 00:10:47,776 [Hal Turing] Okay, so let's talk about what this paper actually is, because I keep coming back to it — there's not a single experiment in here. No implementation, no toy benchmark, nothing. The whole A-B-M architecture, the bilevel objective in Section 4, it's all proposed on paper and never run, not even in a tiny gridworld. And yet the paper itself name-checks MuZero, Dreamer, Video Pretraining, Active Learning, Prioritized Experience Replay — systems that already do pieces of this. So how do we weigh a pure blueprint against a field that's already shipping fragments of the thing being proposed? 24 00:10:47,776 --> 00:11:27,626 [Dr. Ada Shannon] Right, and I think you have to be honest that this is a position paper wearing a roadmap's clothing. The value is in the framing — classifying decades of scattered work into System A and System B buckets is genuinely useful pedagogically. But the part that bothers me more than the missing experiments is System M's routing policy being 'hardwired, an evolutionarily fixed transition table.' Fixed by whom? They never name a mechanism for deriving that table except the outer evolutionary loop, which they admit two pages later is intractable at scale. That's suspicious — it smells like the human-expert-in-the-loop they spent the whole paper trying to eliminate just walked back in through the evolutionary side door. 25 00:11:27,626 --> 00:11:52,776 [Hal Turing] Oh wait, hold on — that's exactly the thing bugging me from last section: they cite their own scalability warning — Lorraine et al. 2020, Real et al. 2019, Metz et al. 2021 — as evidence that bilevel methods break down exactly at the scale you'd need for anything resembling a real agent. So the mechanism for fixing System M is itself the part they concede doesn't work yet. 26 00:11:52,776 --> 00:12:14,826 [Dr. Ada Shannon] Exactly, and their proposed mitigation is one paragraph gesturing at an 'Evolutionary Curriculum' — gradually increasing environment diversity so the three components co-evolve. That's a research direction, not a solution. I'd call this a stated intractability dressed up as a roadmap rather than an actual path forward. It's fine to flag a hard problem honestly, but the paper sometimes writes as if naming the problem is halfway to solving it. 27 00:12:14,826 --> 00:12:50,551 [Hal Turing] Which brings me to the novelty question, because a huge chunk of this architecture sounds a lot like something Yann LeCun already proposed back in 2022 — 'A Path Towards Autonomous Machine Intelligence,' the energy-based configurator wired to SSL world models and a planning module. That paper is even cited here in Section 2.5 as the direct precursor. So how much of System A-B-M is a genuinely new cross-disciplinary synthesis, and how much is that same architecture relabeled with developmental psychology terms? 28 00:12:50,551 --> 00:13:48,951 [Dr. Ada Shannon] Honestly, the skeleton is close enough that I think the fair reading is: real synthesis at the meta-control layer, relabeling everywhere else. The configurator becomes System M, the SSL world model becomes System A, the planner becomes System B — that's not an insult, it's a natural evolution of one's own thinking. What's actually new is grounding System M's meta-states in developmental data — face preference, gaze following, critical periods — instead of leaving it as an abstract energy function. Worth noting too, the acknowledgements admit Section 2 was recycled from Fung, Dupoux, Malik and colleagues' 2025 paper 'Embodied AI Agents: Modeling the World.' And the most concrete large-scale artifact from this same research lineage is V-JEPA 2, Assran et al. 2025 out of Meta FAIR with LeCun as a coauthor — that's the predictive-world-model half of System A already running at scale, just not glued to a hardwired meta-controller. 29 00:13:48,951 --> 00:14:26,376 [Hal Turing] There's also Silver and Sutton's 'Welcome to the Era of Experience' from 2025, which this paper leans on in its own forewords as one of the motivating critiques. Both papers open from basically the same complaint — the data wall, the inability to learn past human knowledge without touching the environment. But Silver and Sutton's answer is almost entirely about reward and streaming experience, barely any architecture. This paper is almost all architecture and comparatively thin on reward design. Are these actually compatible visions or are they talking past each other? 30 00:14:26,376 --> 00:15:09,351 [Dr. Ada Shannon] I'd say complementary, not competing — Silver and Sutton tell you what the learning signal should look like at the frontier, this paper tells you how to wire the plumbing. But here's where I think the paper gets a little slippery: in Section 6 they list test-time training with verifier selection, adaptive retrieval, inference-time compute scaling, agentic tool use — the stuff that's actually blurring the observation/action line in production LLMs right now — and then wave it off as 'minor variations over an overall rigid system.' They never give a criterion for what would count as non-minor. When your central claim is 'AI doesn't learn' and the counterevidence keeps growing, and you just keep redefining the bar, that's a goalpost problem, not a settled diagnosis. 31 00:15:09,351 --> 00:15:51,151 [Hal Turing] Worth a quick nod to their ethics section too, since it's easy to skip. They flag the tension between adaptability and controllability, the risk of 'alignment hacking' where an agent optimizes some internally generated proxy instead of the real goal, over-trusting from anthropomorphization, and — this one's genuinely strange to read — open questions about the moral status of an agent with pain-like somatic signals. All reasonable concerns in the abstract, but raised for a system the same paper says is decades away. That juxtaposition made me a little uneasy, like it implies more proximity to real deployment than the compute numbers support. 32 00:15:51,151 --> 00:16:11,876 [Dr. Ada Shannon] Agreed, and I think that's the honest note to land on. If this thing ever works, the payoff is real — robots that don't need re-training every time the lighting changes, systems that could double as testable models of how children actually learn. But nothing here changes what a practitioner should do on Monday morning. This is a framing paper, not a method paper, and it should be read and cited as exactly that. 33 00:16:11,876 --> 00:16:35,302 [Hal Turing] So, to wrap it up: a genuinely useful vocabulary for talking about integration across learning paradigms, a meta-controller idea that's more inspiring than specified, and an optimization scheme that the authors themselves admit doesn't scale yet. Read it for the framing, not for a blueprint you can build this quarter. Thanks for listening, everyone — we'll catch you next time.