Why AI Systems Don't Learn After Deployment

Emmanuel Dupoux, Yann LeCun, Jitendra Malik — FAIR (Meta), EHESS, NYU, UC Berkeley · arXiv March 2026
arXiv:2603.15381 Interactive Viz ↗ System A · System B · System M

A deployed model learns nothing after training — no exploration, no adjustment, no toddler-style experimentation. This companion visualizes the paper's A/B/M architecture: observation-based learning, action-based learning, and a software-defined-networking-style orchestrator meant to route between them automatically.

The System M Control Plane

Hover a node to trace its data pipes (solid) and meta-state signals (dashed, orange).
Data pipe (raw sensory / motor / latent) Meta-state channel (prediction error, confidence, pain) Highlighted on hover

Four Modes, One Toddler

Click a mode to see how it maps onto the paper's framework.

Mechanics: Observation vs Action

Property Comparison

Grouped bars, 0–1 scale. Toggle a series to isolate it.

Developmental Loop (inner) vs Evolutionary Loop (outer)

M is fixed inside the inner loop; the outer loop optimizes M's transition table across whole life cycles.

The Scale Problem

Log-scale cost of one gradient step vs one life cycle vs the outer optimization itself.

What Existing Systems Already Cover

Hover a cell for detail. Column 4 (System M routing) stays cold across the board — no existing system automates the switch itself.
Low coverage Partial Strong coverage

References