1 00:00:01,000 --> 00:00:49,750 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into "Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems," by Jiacheng Liu et al. — four authors total, including Xiaohan Zhao, Xinyi Shang, and corresponding author Zhiqiang Shen — out of VILA Lab at the Mohamed bin Zayed University of Artificial Intelligence, with a co-author also affiliated with University College London. It went up on arXiv on July 2nd, 2026. Right out of the gate there's a number worth sitting with: the actual model-reasoning core of Claude Code is basically one while-loop. Everything else — permissions, context management, tools, subagents — is a mountain of scaffolding built around it. 2 00:00:49,750 --> 00:01:23,975 [Dr. Ada Shannon] Right, and that's not a throwaway detail, Hal — that's the paper's whole north star. If the model itself just decides what to do next, how much architecture do you need around it to make that trustworthy, resumable, and safe in a tool millions of developers now touch daily? So they go source-diving into the actual extracted TypeScript codebase — version 2.1.88 of Claude Code — and hold it up against two other agent systems, OpenClaw and Hermes Agent, to see whether the same design questions get answered the same way everywhere. Spoiler: they don't. 3 00:01:23,975 --> 00:01:50,425 [Hal Turing] Wait, hold on — Claude Code isn't open source, Anthropic ships it as a compiled CLI, not a public repo. So how do you 'source-analyze' something like that? Anthropic publishes user-facing docs, and there's that official line about an 'agentic loop' that plans and executes toward a goal, but that's product copy, not an architecture spec. Are we talking about someone actually extracting the real TypeScript, or is this inference from watching the tool behave? 4 00:01:50,425 --> 00:02:27,325 [Dr. Ada Shannon] It's the real thing, not guesswork. The npm-distributed package ships a TypeScript build that can be extracted and read — that's how the authors got specific file names like query.ts and permissions.ts into a paper at all. It's literal source, cross-referenced against Anthropic's own docs and outside community analysis. Running it against OpenClaw, a multi-channel personal-assistant gateway, and Hermes Agent from Nous Research, a single-process, multi-surface assistant, gives them a control group — same recurring design questions, three different deployment contexts, three different answers. 5 00:02:27,325 --> 00:02:40,900 [Hal Turing] Okay, that reframes the whole paper for me — this isn't just 'here's how Claude Code works,' it's 'here's what its creators believe matters, made visible in code.' So what are these five human values they pull out? 6 00:02:40,900 --> 00:03:11,300 [Dr. Ada Shannon] Five of them, and they're not fluffy mission-statement stuff — each maps to actual code. Human decision authority: the user keeps ultimate say over what the system does. Safety, security, and privacy: protecting people and their code even when they're not paying close attention. Reliable execution: does the agent do what you actually meant, and stay coherent across a long session instead of drifting. Capability amplification: does it let you do genuinely new things, not just type faster. Contextual adaptability: does the relationship between you and the tool actually improve over time. 7 00:03:11,300 --> 00:03:31,650 [Hal Turing] So let's ground this in the actual machinery then, because 'agentic loop' gets thrown around constantly and I don't think everyone actually knows what's inside that box. And while you're at it — what do people mean by 'deny-first permissions' and this 'compaction pipeline' I keep seeing referenced? Start wherever makes sense. 8 00:03:31,650 --> 00:04:12,500 [Dr. Ada Shannon] Sure. The agentic loop is exactly what it sounds like — a while-loop, literally called queryLoop() in the source, that calls the model, and if the response comes back with a tool-use request, the harness — the code around the model, not the model itself — dispatches it, gets a result, and feeds that back in for another round, until the model answers with plain text and no more tool calls. Deny-first permissions is the safety rule inside that loop: any deny rule overrides every allow rule, even a more specific one — nothing sneaks through on a technicality. And the compaction pipeline handles the fact that context windows aren't infinite: before every model call, five sequential stages shrink the conversation history, each cheaper than the next, so the loop doesn't choke on its own memory. 9 00:04:12,500 --> 00:04:23,926 [Hal Turing] Oh wait, hold on—that context thing — isn't that the same protocol that just got handed over to the Linux Foundation? MCP? Or am I mixing up two different things here? 10 00:04:23,926 --> 00:05:20,201 [Dr. Ada Shannon] You're not wrong, but you're mixing two ideas together. MCP, the Model Context Protocol, is the standardized way Claude Code talks to external tool servers — databases, ticketing systems, whatever a third party wants to expose — and yes, it was donated to the Linux Foundation's Agentic AI Foundation, so it's no longer purely Anthropic-controlled. What you're actually thinking of is subagent delegation: the main agent spawns a child instance through the harness, and that subagent gets its own isolated context window, does its work, and reports back just a summary, not its whole transcript. Zoom out and the whole system is seven components — user, interfaces, the loop, permissions, tools, state, and the execution environment — layered into five subsystems underneath. Right — and that's the outward protocol. What we haven't touched is the plumbing on this side of it — what happens between the model deciding to call a tool and that tool actually running. There's a lot more machinery in there than 'loop calls model, loop calls tool.' 11 00:05:20,201 --> 00:05:30,226 [Hal Turing] Okay, walk me through it then. Not the philosophy — the actual mechanics. What happens in the gap between me hitting enter and Claude touching a file? 12 00:05:30,226 --> 00:06:02,501 [Dr. Ada Shannon] Every turn runs a fixed sequence. Settings resolution first — the loop pulls in the immutable stuff, system prompt, model config, permission callback. Then it sets up one mutable State object for everything else — messages, compaction tracking, recovery counters — and there are seven points in the loop, called continue sites, where instead of tweaking individual fields, the whole object just gets reassigned in one shot. Context assembly grabs messages forward from the last compact boundary. Then, before the model is even called, five separate shapers run on that message array to manage— 13 00:06:02,501 --> 00:06:08,651 [Hal Turing] Wait, wait, hold on — five shapers, every single call? Not just when things get long? 14 00:06:08,651 --> 00:06:28,851 [Dr. Ada Shannon] Every call, yeah, they're just cheap when there's nothing to do. After that the model streams back its response, and any tool_use blocks get dispatched, pass through a permission gate, and results flow back in as tool_result messages. The loop keeps going until the model returns plain text with no tool calls — that's the real stop condition, not a counter, though there's a max-turns cap sitting behind it as a backstop. 15 00:06:28,851 --> 00:06:39,351 [Hal Turing] So that permission gate — what's actually deciding allow versus deny in there? Because 'deny-first' is a philosophy, but there's got to be a mechanism. 16 00:06:39,351 --> 00:07:17,101 [Dr. Ada Shannon] Building on deny-first — blanket-denied tools get stripped from the model's view entirely, so it can't even try. Then there are seven permission modes — plan, default, acceptEdits, auto, dontAsk, bypassPermissions, and an internal-only one called bubble that escalates a subagent's request up to the parent terminal. In auto mode there's an ML classifier doing the judgment call. And that classifier exists because of a number Anthropic published — users approve about ninety-three percent of permission prompts, Hughes 2026. Once approval becomes a reflex, interactive confirmation alone can't be the safety mechanism anymore. 17 00:07:17,101 --> 00:07:23,676 [Hal Turing] And the five shapers — what are they actually doing, and why five instead of just one summarizer? 18 00:07:23,676 --> 00:07:58,626 [Dr. Ada Shannon] Because no single strategy handles every kind of pressure. Budget reduction caps individual tool outputs. Snip trims older history. Microcompact does cache-aware fine compression. Context collapse is a read-time projection — it doesn't even touch the stored history, just what the model sees. Auto-compact is the expensive one, a full model-generated summary, and it only fires if the first four didn't free enough room. Cheapest first, escalate only when needed. That's the same instinct behind assembling up to fifty-four tools and twenty-seven hook events — most of it stays dormant until a specific situation actually calls for it. 19 00:07:58,626 --> 00:08:08,151 [Hal Turing] Okay, and when Claude decides it needs help — delegates part of the auth-test job to a subagent — what does that actually look like under the hood? 20 00:08:08,151 --> 00:08:41,177 [Dr. Ada Shannon] The Agent tool spins up a new queryLoop() instance for it. It runs, does its exploring or editing, and the full subagent transcript gets written to its own sidechain file, never entering the parent's context — only that final summary comes back, like we said. Schluntz and Zhang's 2024 Anthropic writeup calls this the orchestrator-workers pattern, and it's deliberate — conversation-sharing frameworks that pass full histories between agents risk exploding the context budget as the agent count grows. 21 00:08:41,177 --> 00:08:48,152 [Hal Turing] Given all that, how does this actually compare to OpenClaw and Hermes now that we've seen the internals? 22 00:08:48,152 --> 00:09:48,027 [Dr. Ada Shannon] Pretty different bets. OpenClaw, that's Steinberger and the OpenClaw contributors, 2026, wraps the whole thing in a gateway control plane and puts the trust boundary at the perimeter — one trusted operator, DM pairing, allowlists, sandboxing that's opt-in rather than default. Hermes Agent, Nous Research 2026, is a single process, but keeps per-action approval like Claude Code, rendering the same approval flow across four different surfaces — CLI, Telegram, Discord keyboards, an ACP round-trip for IDEs. And none of this is static, by the way — McCain and colleagues, 2026, found auto-approve rates climb from around twenty percent in a user's first fifty sessions to over forty percent by session seven-fifty. The trust boundary itself moves. So you've got three genuinely different trust models: Claude Code's per-action gate sitting inside the loop, OpenClaw's perimeter checkpoint sitting outside it, and Hermes doing per-action checks but without Claude Code's subagent isolation. None of them agrees on where the trust boundary should live, which tells you the field hasn't converged on an answer. 23 00:09:48,027 --> 00:10:08,302 [Hal Turing] Okay, now that we've seen all three side by side, I want to push on something that bugged me reading this. That 1.6 versus 98.4 percent number, decision logic versus operational infrastructure, is basically the paper's whole thesis in one stat. Where does that number actually come from? 24 00:10:08,302 --> 00:10:49,402 [Dr. Ada Shannon] That's exactly the problem. They define their own three-tier evidence system: Tier A is official Anthropic docs, Tier B is stuff they directly verified against the extracted source, Tier C is community analysis and inference, stated with hedging language by their own rule. The 1.6/98.4 split is Tier C. It's not something the authors measured by walking the codebase and classifying every file themselves, it's borrowed from someone else's estimate. And it gets repeated three separate times as the anchor for their entire minimal-scaffolding argument. A paper careful enough to build a three-tier system and then hang its single most-quoted number on the tier it flagged as weakest, that's a load-bearing wall built on their own least-reliable evidence. 25 00:10:49,402 --> 00:11:10,752 [Hal Turing] Wait, hold on, that's actually worse than it sounds, because this whole analysis is one snapshot, right, v2.1.88. And they admit in their own limitations section that feature flags can make different builds functionally different applications. So even if that 1.6 percent number were solid, solid for which build? 26 00:11:10,752 --> 00:12:03,277 [Dr. Ada Shannon] Right, and that cuts deep. Auto-mode classifier, context collapse, cached microcompact, three of the flashiest mechanisms in the whole paper are gated behind feature flags in their own source map. If TRANSCRIPT_CLASSIFIER isn't flipped on in the build you're running, there's no classifier evaluating your tool calls, full stop. They're honest about that risk in the limitations section, but the main text doesn't carry the caveat forward, it describes these as the architecture, present tense, when what they've got is one commit's worth of possible architecture. There's a layer under that too: how did they get this source at all? A publicly available npm package extraction of closed-source software, with their own disclaimer that the IP belongs to Anthropic. That's not a neutral footnote, it's an open question about whether this was theirs to publish in this much detail, and it should temper how much weight you put on their ground-truth claims. 27 00:12:03,277 --> 00:12:27,477 [Hal Turing] And that provenance question echoes into how they use their own security findings. They cite CVE-2025-59536 and CVE-2026-21852, the pre-trust hook and MCP execution bugs, as evidence the layered-defense design works. But those are cases where the defense failed and someone got through, isn't that backwards? 28 00:12:27,477 --> 00:13:24,152 [Dr. Ada Shannon] It's exactly backwards. Presenting successful exploits as validation treats eventually-patched as proof the architecture was sound, when what actually happened is there was a real gap during that pre-trust window. And here's what's missing entirely: they describe that vulnerability class, hooks and MCP servers executing before the trust dialog resolves, without once citing Greshake, Abdelnabi, Mishra, Endres, Holz, and Fritz, Not What You've Signed Up For, out of CISPA Helmholtz Center for Information Security, 2023. That's the foundational indirect prompt injection paper, it's literally the threat model this permission and hook system defends against, and it's absent. Same gap on evaluation: it's all SWE-Bench and HumanEval, narrow coding benchmarks, zero engagement with GAIA, Mialon and colleagues out of Meta AI and Hugging Face, 2023, which tests general-purpose agent competence. So we don't actually know if this harness design generalizes past fixing a failing test. 29 00:13:24,152 --> 00:13:37,802 [Hal Turing] So if I'm someone actually building or evaluating an agent harness right now, what's the takeaway, invest in the deterministic infrastructure the way they argue, but hold the specific numbers loosely? 30 00:13:37,802 --> 00:14:21,952 [Dr. Ada Shannon] Pretty much. The structural argument, that a rich deterministic harness lets the model's judgment do more work than a rigid planner would, holds up independent of the shaky number, because you can watch that mechanism directly in verified Tier B source. What holds up less well is anything self-reported through Anthropic's own surveys: the 93 percent approval rate from Hughes 2026, the 27 percent of tasks that wouldn't have been attempted from Huang and colleagues' internal survey of 132 engineers. Anthropic has a direct commercial interest in that story landing exactly this way, safe enough to justify more autonomy, transformative enough to justify the subscription. Doesn't make the numbers false, but it's marketing-adjacent evidence, not independent measurement. 31 00:14:21,952 --> 00:15:03,127 [Hal Turing] So coming back to where this whole conversation started, how Claude Code's architecture answers these design questions and whether it generalizes: the mechanism itself is real and genuinely well documented at the code level, one loop, layered permissions, five-stage compaction, isolated subagents. But the paper's biggest claims, the ratio anchoring its whole philosophy, the safety story built on exploited CVEs, a comparison against just two other systems, are running ahead of what one reverse-engineered snapshot can prove. Good architecture, overconfident framing. That's it for this one, thanks for listening.