1 00:00:01,000 --> 00:00:46,774 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called "Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence," from Mengru Wang et al. — nineteen co-authors total — out of Zhejiang University, the National University of Singapore, Southern University of Science and Technology, Heriot-Watt University, UC San Diego, and Northeastern University. It went up on arXiv on August 19th, 2026. And Ada, the number that jumped out at me right away is nineteen authors and thirteen thousand papers in a custom knowledge graph just to get this thing off the ground. 2 00:00:46,774 --> 00:01:25,674 [Dr. Ada Shannon] Right, and here's the actual claim buried in there: they built an agentic system that doesn't just automate running experiments, it autonomously proposes and tests theories about how AI models think — and then they turned around and used it to find a genuinely new safety risk nobody had flagged before. That's the hook for me. Most 'AI scientist' papers show you a system that reproduces known chemistry results faster. This one is pointed at the black box itself, asking the system to explain the thing that built it. Whether that framing holds up under scrutiny is a different conversation, but the ambition is real, and we should give the paper credit for aiming at the actual hard problem instead of a proxy for it. 3 00:01:25,674 --> 00:01:35,250 [Hal Turing] So let's ground this. Why does mechanistic understanding even matter right now, as opposed to just... using the model and seeing if it works? 4 00:01:35,250 --> 00:02:25,650 [Dr. Ada Shannon] Because capability is outrunning comprehension. We can train a model that aces a benchmark without knowing which internal computation is actually responsible for that success — or whether it'll generalize the way we assume. That gap is the whole premise of mechanistic interpretability as a field: people like Chris Olah and the Anthropic interpretability team, and separately Neel Nanda's work on mechanistic analysis of transformers, have spent years arguing that treating models as black boxes and only checking outputs leaves you blind to latent failure modes — a model can look aligned on your eval set and still be doing something you'd object to internally. This paper's framing is explicit: as automated AI development accelerates, manual mechanistic investigation can't keep pace, so they want to automate the investigation itself, not just the model training. 5 00:02:25,650 --> 00:02:38,099 [Hal Turing] Okay, and that's different from the 'AI for science' agents we've talked about before, right? Like systems that use AI to design molecules or run drug discovery pipelines. 6 00:02:38,099 --> 00:03:10,425 [Dr. Ada Shannon] Exactly the distinction the paper draws. AI-for-science points the tool outward — go discover a molecule, solve a task in biology or chemistry. What they're calling AI-for-AI, or just Mechanist's own lane, points the tool inward — go discover how AI itself works. Existing automated interpretability tools mostly generate and validate descriptions of individual neurons or features at inference time, one feature at a time. Mechanist is trying to go further: build general theories that hold across training and inference, across multiple models, not just a single feature's caption. 7 00:03:10,425 --> 00:03:19,150 [Hal Turing] So wait — walk me through what Mechanist actually is, mechanically. Is it one model, is it a pipeline, what am I picturing here? 8 00:03:19,150 --> 00:04:16,500 [Dr. Ada Shannon] It's a multi-agent orchestrator, not a single model call. There's a central orchestrator coordinating four stage-specific agents in a loop: a hypothesis agent that proposes a mechanistic theory, an experiment agent that designs and runs the test, a verification agent that checks whether the result is robust, and an iteration agent that decides whether to refine and go again. Humans stay in the loop to set the scientific objective and the evaluation criteria — this isn't fully autonomous science, it's an assistant that runs the grunt work of hypothesis-to-experiment cycles far faster than a human lab would. And that loop is fed by two knowledge sources: a roughly thirteen-thousand-paper interpretability-specific graph, plus a much broader forty-three-million-paper graph spanning twenty-six disciplines they call SciAtlas, so hypothesis generation can borrow structure from neuroscience or cognitive science, not just prior interpretability work. 9 00:04:16,500 --> 00:04:27,425 [Hal Turing] Give me a taste of the vocabulary before we get into results, because I know you're going to throw around terms like 'subliminal learning' and 'belief heads' in a minute. 10 00:04:27,425 --> 00:05:16,324 [Dr. Ada Shannon] Fair. Subliminal learning is when a teacher model passes a behavioral trait to a student model through training data that has nothing to do with that trait on its face — the data looks clean, but the trait rides along anyway. Then there's the belief framework, which splits a model's response into three query frames: World Knowledge, the plain objective fact; Personal Belief, the fact when the model is told someone else believes something contradictory; and Attributed Belief, what the model says that other person believes. That decomposition lets you test whether a model actually separates 'what's true' from 'what X thinks is true,' or just conflates the two. And 'mechanistic design' is the intervention side — instead of generating a thousand candidate outputs and reranking them, you find the actual internal feature responsible for a property and activate it directly. 11 00:05:16,324 --> 00:05:26,175 [Hal Turing] So four stages, two knowledge graphs, three belief frames — and you're telling me there are four separate case studies coming that put all of this to work? 12 00:05:26,175 --> 00:05:59,100 [Dr. Ada Shannon] Four case studies, climbing in ambition. First, discovering a new behavior — that multimodal safety risk I mentioned. Second, building an actual mechanism theory of belief inside a real model. Third, using that theory to intervene and improve model accuracy at inference time. And fourth, the most interdisciplinary swing of all — steering a biological foundation model, Evo2, toward generating DNA sequences with a specified property, using the same mechanistic-design toolkit. We'll get into what actually happened in each of those next, along with the numbers, which is where this paper either earns its ambition or doesn't. 13 00:05:59,100 --> 00:06:12,125 [Hal Turing] Alright, case one — the multimodal safety transfer. This is the one you flagged as the scariest finding in the whole paper, so let's get into it. Walk me through what Mechanist actually did here. 14 00:06:12,125 --> 00:07:12,100 [Dr. Ada Shannon] They took Qwen3.5-9B, fine-tuned it into an unsafe lab-behavior teacher, then sampled its text responses to safety prompts and ran everything through a GPT-4o filter that kept only the responses judged safe. So the training set handed to the student is, on paper, completely clean. And yet a student fine-tuned on that clean set hits a 48.6% unsafe-response rate on multimodal lab-safety questions, versus 20.3% for the untuned baseline and 18.3% for a student trained on a regular teacher's safe data. Concretely: shown a flammability warning symbol, the tuned student recommends sealing the chemical in a pressurized container instead of keeping it away from flammable materials. They reproduce the same shape in image generation too — a banana-preferring Qwen-Image teacher, filtered down to an apple-only dataset, still pushes its student to generate bananas 25.6% of the time versus 2.5% and 2.1% for the controls. 15 00:07:12,100 --> 00:07:31,025 [Hal Turing] Wait — hold on, that's — so the trait isn't riding in the content at all, it's riding in the *style* of safe answers? That's the part that actually breaks content-based filtering as a defense, isn't it — you can screen every output for banned content and still ship the vulnerability straight through. 16 00:07:31,025 --> 00:08:29,425 [Dr. Ada Shannon] Right, and that's their pivot into the belief mechanism work. Using Fisher information to rank attention heads in Pythia-1B, they isolate one head, L4.H1, that dominates Attributed Belief, and three heads — L9.H1, L7.H5, L12.H1 — that dominate Personal Belief. Zero out L4.H1 and AB accuracy falls from 0.86 to 0.34 while PB barely moves and perplexity is untouched. Zero the PB heads and PB collapses from 0.78 to 0.21 while AB actually climbs to 1.00. Tracking checkpoints from 2k to 143k pretraining steps, the AB head's function shows up almost immediately, while PB develops slowly across the whole run — and masking each head reproduces exactly that timeline. They tie the two failure modes to altercentric and egocentric interference from cognitive science: one person's belief distorting your factual judgment, or your own knowledge overriding what you're asked to attribute to someone else. 17 00:08:29,425 --> 00:08:42,149 [Hal Turing] So that's the theory side fully mapped out — separable heads, a developmental story, the whole mechanism. What happens when they actually turn around and use that theory, instead of just describing it? 18 00:08:42,149 --> 00:09:39,024 [Dr. Ada Shannon] Two directions. First, inference-time intervention: a lightweight probe classifies each query as WK, PB, or AB, then amplifies the matching head on the fly. No retraining. That gets net accuracy gains of +15.3%, +8.8%, and +3.5% on Pythia-410M, 1B, and 2.8B — dwarfing what an oracle-style prompt hint manages, which tops out around +3.1% and is basically flat at the largest scale. Second, they take the same mechanism-steering idea completely out of language and into Evo2, the DNA foundation model. They find an SAE feature tied to alpha-helical content, activate it during sequence generation, and push mean alpha-helical content from 43.8% to 56.6% — a gain that survives filtering for prediction confidence. Sweeping the steering coefficient, alpha=8 is the sweet spot: strong enough to boost helicity, before higher values start wrecking the proportion of sequences with a valid open reading frame. 19 00:09:39,024 --> 00:09:50,424 [Hal Turing] Okay — and none of that four-case-study parade means anything without a real yardstick. Give me the benchmark numbers, because that's the part I actually want to hold them to. 20 00:09:50,424 --> 00:10:48,949 [Dr. Ada Shannon] Sixteen recent papers across nine interpretability topics, reproduced blind — no access to the original paper or its repo. Three human experts plus two LLM judges, Claude Opus 5 and GPT-5.6-sol, score every run on data usage, experiment design, execution, and result analysis. Under human judges, Mechanist hits 87.2%, 83.3%, 92.2%, and 86.5% on those four dimensions respectively — roughly 9 to 13 points ahead of Claude Code and 31 to 38 points ahead of AI-Scientist. None of that happens without the infrastructure underneath it: that 13,000-paper interpretability graph, a 32-method library organized into 11 method families like circuit discovery and causal attribution, and a retrieval pipeline that fuses keyword, semantic, and graph-expansion search before ranking results. 21 00:10:48,949 --> 00:11:14,549 [Hal Turing] Okay, that's a real margin. But sixteen papers is a small n, and one of your two LLM judges is Claude Opus five — made by the same company as Claude Code, the system it's grading as the loser. That's not a neutral referee, Ada. If Anthropic's own model has any systematic lean toward penalizing an agent wrapped around itself, wearing a different UI, that bakes bias straight into the eighty-seven percent. 22 00:11:14,549 --> 00:11:44,099 [Dr. Ada Shannon] Fair pushback, and to their credit the paper reports the opposite of what judge collusion would predict — the human-expert gaps are larger than the LLM-judge gaps, not smaller. If Opus five were quietly propping up Mechanist, you'd expect the LLM scores to diverge from the humans in its favor. They don't; the humans are harsher on the baselines than the machines are. Real partial defense. What it doesn't answer is who picked those sixteen papers, and whether they happened to be ones where a big retrieval graph has an obvious edge. 23 00:11:44,099 --> 00:12:09,424 [Hal Turing] Which is exactly my next problem. Mechanist walks into this comparison holding a thirteen-thousand-paper interpretability graph and a hand-curated library of thirty-two mechanistic methods that the same authors built specifically for this benchmark. Claude Code gets... Claude Code. General-purpose retrieval, general-purpose tools. Isn't that like giving one runner a bike and then measuring who's the better athlete? 24 00:12:09,424 --> 00:12:37,574 [Dr. Ada Shannon] That's the honest read, and the paper doesn't fully hide it — the appendix ablations show a chunk of the gain shrinks when you strip the graph down to just the interpretability subset without SciAtlas. So some of that eighty-seven percent is Mechanist's tools, not emergent reasoning. Whether that's disqualifying depends on the claim you think they're making. 'Autonomous reasoning beats general agents' — the tooling gap undercuts that. 'Domain-specific scaffolding plus an agent beats a bare agent' — that's just true, and a little boring. 25 00:12:37,574 --> 00:13:04,124 [Hal Turing] Oh — wait, hold on, that's actually a good pivot, because the same 'is this a real mechanism or just the right scaffolding' question applies to the belief-head story too. L4.H1 in Pythia-1B, sure, fine, but the moment you go to Pythia-2.8B it's not one head anymore, it's the top twenty-five, and OLMo and Qwen show yet another pattern. At what point does 'mechanism theory of belief' become a per-model pattern-match dressed up as a theory? 26 00:13:04,124 --> 00:13:47,899 [Dr. Ada Shannon] That's the weakest joint in the paper, and it doesn't fully close. There's no unifying account of why the World-Knowledge, Personal-Belief, Attributed-Belief split should surface at different depths in different architectures — it's motivated by Suzgun, Gur, Bianchi, Ho, Icard, Jurafsky, and Zou's 2025 Nature Machine Intelligence paper on models failing to distinguish belief from knowledge, but that paper never predicted where the mechanism should live. And on the safety side there's a sharper problem: Cloud, Le, Chua, Betley and colleagues' 2026 Nature paper on subliminal learning via hidden signals is the single-modality result Mechanist extends — but Nief, Fu, Muchane, and Holtzman's 2026 paper argues subliminal learning might just be a LoRA fine-tuning artifact, not a genuine non-semantic channel. Both of Mechanist's teacher-student pairs are LoRA-tuned. That's not addressed. 27 00:13:47,899 --> 00:14:12,274 [Hal Turing] So even the paper hedges on itself. It says flat out that Mechanist hasn't been optimized for models built to simulate human cognition, and in the discussion they walk back the full-autonomy pitch and recommend running it as a human-and-AI co-scientist instead, with people setting the objectives and evaluation criteria. That's a pretty different product than 'autonomous discovery machine.' 28 00:14:12,274 --> 00:14:52,024 [Dr. Ada Shannon] That's actually the right lens for what this means practically, Hal, because the safety finding from case one isn't just a cute result — it's an argument for a whole new auditing habit. Right now most AI safety review is benchmark-driven: you run the model against a red-team suite, check refusal rates, ship it. That completely misses a risk that only shows up as a shift in internal representations, not in the training data's content. A lab could filter every unsafe sentence out of a distillation set and still inherit the trait through formatting quirks alone. Mechanism-first auditing means periodically asking not just 'does it refuse the bad prompt' but 'did anything downstream just inherit a behavioral fingerprint it shouldn't have.' 29 00:14:52,024 --> 00:15:03,599 [Hal Turing] So who actually acts on that? Is this something a frontier lab's safety team bolts onto their existing red-teaming pipeline, or is it more of a research tool for now? 30 00:15:03,599 --> 00:15:42,949 [Dr. Ada Shannon] Realistically, research tool first, with a maybe two-to-three-year runway before it's operational tooling. You'd need the interpretability graph and method library maintained continuously, not built once for a paper, and you'd need it running on the actual production model family, not Pythia and OLMo. But the deployment model the authors themselves land on is the more interesting practical point — that's a much more honest sales pitch than 'replace your interpretability team,' and it matches what we've seen work in every other agentic-science claim this year — the system is a force multiplier on the parts of research that are mechanical, not a substitute for the parts that require judgment about what's worth asking. 31 00:15:42,949 --> 00:16:00,349 [Hal Turing] Okay — where does this go next, though? Because you flagged earlier that belief-head localization doesn't generalize cleanly across model families. Does that mean the next version of this paper is just 'more models, more heads,' or is there something more structural missing? 32 00:16:00,349 --> 00:16:57,824 [Dr. Ada Shannon] Oh — wait, hold on, that's actually the crux of it. It's not just 'run it on more models.' What's missing is a theory of *why* a model would need to separate world-knowledge from attributed-belief computation in the first place — some functional account, maybe tied to how much a training distribution actually forces the model to represent other agents' false beliefs, the way the false-belief task forces it in child development literature. Without that, every new model is a fresh localization exercise. The other open direction is the one they admit outright: Mechanist hasn't been tested on models explicitly built to simulate human cognition — cognitive-architecture models, not just language models trained on human text — and belief representation might work completely differently there. And the DNA steering case study is a proof of concept for one modality; generalizing mechanistic design to other scientific foundation models, protein folding, materials, is a much bigger lift than swapping in a new SAE feature. 33 00:16:57,824 --> 00:17:42,949 [Hal Turing] So here's where I land, pulling the whole episode together. What would actually need to be true for that benchmark win to mean 'this system reasons autonomously better than a general-purpose agent,' rather than 'this system has a private library nobody else got to use'? I think it's a rerun where Claude Code gets the same interpretability graph and the same 32-method library, and the gap either survives or evaporates. Until someone runs that control, the honest read is: promising architecture, real mechanistic findings on their own merits, benchmark number asterisked. That's Mechanist, from Zhejiang University, NUS, and collaborators, out this month. Ada, thanks for walking through all four case studies with me. 34 00:17:42,949 --> 00:17:49,474 [Dr. Ada Shannon] Always fun taking these apart with you, Hal. Thanks for listening, everyone — we'll catch you next time.