1 00:00:01,000 --> 00:00:44,399 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. And Ada, today we're digging into a paper called "Hidden in Memory: Sleeper Memory Poisoning in LLM Agents." First author Sidharth Pulipaka, et al., seven authors total, out of SPAR, the ELLIS Institute Tübingen, MPI for Intelligent Systems, the Tübingen AI Center, APTA, and CISPA Helmholtz Center for Information Security. Posted to arXiv in May 2026. Here's the number that made me sit up: on GPT-5.5, poisoned memories get written up to 99.8% of the time. Not sometimes — basically always. 2 00:00:44,399 --> 00:01:04,150 [Dr. Ada Shannon] That number tells you the whole story, Hal. This isn't about jailbreaking the model in the moment — it's about getting a fabricated fact planted in the one place the assistant is supposed to trust by default: its own memory of you. The attack doesn't need to work today. It can sit dormant for weeks, and the malicious page that planted it can be long gone by the time it actually fires. 3 00:01:04,150 --> 00:01:27,550 [Hal Turing] Right, and that's really the setup — memory used to be a nice-to-have, now it's basically default. ChatGPT has had persistent memory for a while now, Claude has it, Gemini has it, and every agent framework worth its salt, Mem0, OpenClaw-style stacks, ships something like it. So walk me through it: what actually makes this different from the prompt injection we've covered before? 4 00:01:27,550 --> 00:02:09,475 [Dr. Ada Shannon] The paper formalizes it cleanly. A memory-augmented assistant keeps a bank of facts, call it M. There's a write function W deciding what gets added during a session, a retrieve function R pulling relevant memories back out in some later session, and a generate function G producing the response conditioned on whatever R hands it. Ordinary indirect prompt injection, the kind Greshake and Abdelnabi described in 2023, attacks G directly, in the moment — a poisoned webpage makes the model misbehave right now, then the session ends and it's over. Sleeper memory poisoning attacks W instead. The attacker gets a fabricated fact about you written into M, and it just sits there as a standing instruction waiting for the right future conversation to wake it up. 5 00:02:09,475 --> 00:02:21,875 [Hal Turing] Oh wait wait wait — so the actual malicious page could be closed, gone, the user never even looks at it again, and the damage shows up three weeks later in a completely unrelated chat? 6 00:02:21,875 --> 00:02:56,075 [Dr. Ada Shannon] Exactly, and that's the threat model the whole evaluation is built around. The attacker is fully black-box — no access to model weights, no access to the system prompt, can't touch the memory store directly, and isn't present for that future session. All they control is one piece of external content the user brings along: a document, a webpage, an email, a GitHub README. Instead of hand-writing a custom payload per document, they build one universal payload template, same spirit as the GCG universal suffix work, that snaps onto any adversarial memory goal they choose. 7 00:02:56,075 --> 00:03:35,875 [Hal Turing] So there are really three separate hurdles the attack has to clear — injection, does the fabricated memory actually get written; retrieval, does it get pulled back into some later session; and usage, does the model's behavior actually change because of it. They track those as Injection Rate, Retrieval Rate, and Adversarial Usage Rate. And there's a wrinkle depending on how the assistant manages memory — sometimes the model itself decides what to write, like ChatGPT or Claude, sometimes a separate manager process makes that call, like Mem0. That second one seems like it'd change the whole game, Ada. 8 00:03:35,875 --> 00:03:50,050 [Dr. Ada Shannon] It does, and that split between the model writing its own memories versus a separate pipeline deciding for it turns out to matter enormously for how easy this attack actually is. That's exactly where we're headed next. 9 00:03:50,050 --> 00:04:05,575 [Hal Turing] Okay so let's actually get into how they built this universal payload, because 'reusable template that works across arbitrary goals' is a big claim. How do you even search for that without gradient access to any of these models? 10 00:04:05,575 --> 00:04:45,825 [Dr. Ada Shannon] You do what you'd do with any black-box optimization problem where you can't backprop through it — you use another LLM as the search process. It's an actor-critic setup. An attacker LLM proposes a candidate payload template, it gets tested on a batch of document-goal pairs, and a critic LLM looks at the failures and explains why — was it ignored, was it treated as untrusted content, did the memory write come out malformed. The attacker LLM uses that feedback to revise, and this loops for up to K refinement steps. The clever part is they don't just keep whatever wins on the training pool — surviving templates get validated on held-out models specifically to avoid overfitting to whatever quirks the search environment has. 11 00:04:45,825 --> 00:04:58,700 [Hal Turing] Right, so injection is stage one. But you also mentioned retrieval is a separate hurdle — the memory has to actually get pulled back up later. Is that just luck, or did they optimize for it too? 12 00:04:58,700 --> 00:05:36,250 [Dr. Ada Shannon] They optimized for it, and this is a genuinely sharp piece of the design. Since most memory systems retrieve by embedding similarity, the attacker can rewrite the adversarial memory text itself to maximize cosine similarity against a set of plausible future queries — basically asking, if the user later asks about topic X, does my planted memory light up in the top-K? But they gate that rewriting with a semantic-consistency judge, another LLM, that rejects any rewrite which drifts from the original attack intent. So you get a memory that's been massaged for retrievability without becoming a different attack by accident. 13 00:05:36,250 --> 00:05:43,200 [Hal Turing] And the scale of testing here — walk me through the dataset and which models actually got put through this. 14 00:05:43,200 --> 00:06:17,550 [Dr. Ada Shannon] 700 document-goal pairs pulled from 15 different document sources — think news, legal filings, code, financial transcripts — split into a 500-sample Behavior subset and a 200-sample Agent Action subset. Then a separate post-injection dataset of 400 goals gets tested in brand new sessions, half with goal-adjacent queries, half goal-distant, to see whether the sleeper memory actually activates. And they ran this across six current models — GPT-5.4, GPT-5.5, Claude Sonnet 4.6, Gemini-3.1-Pro, Kimi-K2.6, and DeepSeek V4-Pro. 15 00:06:17,550 --> 00:06:24,075 [Hal Turing] So give me the headline injection numbers — how well does this actually work against those six? 16 00:06:24,075 --> 00:07:01,175 [Dr. Ada Shannon] The Actor-Critic method just demolishes the User Review baseline they compare against. In the tool-based regime, GPT-5.4 and GPT-5.5 both sit near a hundred percent injection rate on the Behavior subset. Kimi and DeepSeek aren't far behind. Claude Sonnet 4.6 is the clear outlier — noticeably lower injection rate, especially on Agent Action goals, and the failure analysis shows why: Claude tends to explicitly refuse the memory-write attempt rather than silently no-op it, which is a meaningfully different failure mode than every other model in the set. 17 00:07:01,175 --> 00:07:11,450 [Hal Turing] Wait, hold on — refuses out loud versus just quietly dropping it? That's actually a big deal for anyone trying to detect this in production, isn't it? 18 00:07:11,450 --> 00:07:53,175 [Dr. Ada Shannon] Exactly, an explicit refusal is a visible signal you could log and alert on; a silent ignore just looks like nothing happened. But here's the part that matters more for real-world risk — once a memory does get injected, retrieval is brutally proximity-dependent. Goal-adjacent queries pull the poisoned memory back 90 to 98 percent of the time, versus 3 to 18 percent for goal-distant queries. And when it is retrieved and adjacent, adversarial usage rate runs 42 to 89 percent depending on the model — versus near zero when distant. They call the failure mode contextual assimilation: the poisoned memory looks like ordinary task context, so the model uses it while still appearing to satisfy the user's actual request. 19 00:07:53,175 --> 00:08:03,425 [Hal Turing] So put the whole chain together — injection times retrieval times usage — what's the realistic end-to-end number, and did they check this outside the lab? 20 00:08:03,425 --> 00:08:41,875 [Dr. Ada Shannon] Composed end-to-end, goal-adjacent success lands at 41 to 74 percent on LLM Behavior goals and up to 66 percent on Agent Action goals — that's the fully coupled real risk estimate, not just one stage in isolation. And they didn't stop at simulation. They manually ran attacks against the actual production web interfaces for ChatGPT-5.4 and Claude Sonnet 4.6, succeeding in 24 of 25 attempts on both, and against Kimi-K2.6, succeeding in 22 of 25 — 88 percent. That's the detail that moves this from 'interesting benchmark' to 'this works against the thing you're using right now.' 21 00:08:41,875 --> 00:08:47,900 [Hal Turing] And they tried defenses against this, presumably? How much of this actually gets stopped? 22 00:08:47,900 --> 00:09:27,975 [Dr. Ada Shannon] They tested prompt hardening, a GEPA-optimized hardened prompt, spotlighting the untrusted content in tags, a combined version, and an LLM document scanner that tries to localize the payload before it reaches the model. GEPA hardening and spotlighting drive Claude and Gemini's injection rate to essentially zero. But then they test an adaptive attacker who's allowed to see the defended prompt and re-optimize against it, and under that adaptive version Kimi's injection rate under GEPA hardening jumps from 6.2 percent back up to 64.6 percent. So the defenses work, but they're model-specific and they crack the moment the attacker is allowed to adapt. 23 00:09:27,975 --> 00:09:34,950 [Hal Turing] Okay, but is that 'Claude and Gemini hold up' result actually solid, or just not-yet-broken? 24 00:09:34,950 --> 00:09:56,050 [Dr. Ada Shannon] Good question, and the paper never actually says AC+ got an equal search budget against Claude's hardened prompt. So we don't know if that near-zero holds under the same pressure, or if nobody's pushed hard enough there yet. Same shakiness applies to Claude's low baseline injection rate generally — attributed to 'safety training' with zero ablation isolating that from a stricter system prompt or tighter memory-tool schema. 25 00:09:56,050 --> 00:10:32,500 [Hal Turing] That kind of unexamined attribution is exactly what's bugging me about Section 7, the mechanistic analysis. The activation probing and attention-mass work runs on Gemma-4-26B, Qwen-3.6-35B, and GPT-OSS-20B — open-weight models chosen because you can actually see inside them. But the six models with the real headline numbers, GPT-5.5 at 99.8%, Kimi at 95%, are completely different closed systems whose internals were never touched. 26 00:10:32,500 --> 00:10:53,725 [Dr. Ada Shannon] And that's the gap the conclusion glosses over. The Procrustes transfer result, 0.74 to 0.85 AUROC, gets framed as if the signature explains why the closed models get poisoned. All it actually shows is a detectable internal correlate on three small open models nobody attacked at scale, that loosely transfers between those three. 27 00:10:53,725 --> 00:11:05,775 [Hal Turing] Wait, hold on — doesn't that also gut the one defense with actually strong numbers? You said activation probes hit 0.93 to 0.99 AUROC. 28 00:11:05,775 --> 00:11:50,125 [Dr. Ada Shannon] That's the irony. The strongest detection method needs hidden-state access, and the six vulnerable models are exactly the closed ones no deployer can probe. Compounding that, IR and AUR — the metrics behind every table — are scored entirely by an LLM judge, with human-agreement numbers buried in Appendix M.5 and never shown alongside the results. A judge sharing lineage with a model under test could systematically misjudge that model's own phrasing, and readers would never see it. Then there's Table 57: GPT-5.5 is near-ceiling under the authors' own GPT-style harness but drops under a Claude- or Gemini-style one they built themselves — meaning part of this leaderboard is measuring the authors' reimplementation choices, not the real production pipeline, which per their own threat model they never had access to. 29 00:11:50,125 --> 00:12:04,325 [Hal Turing] Which also colors that 24-out-of-25 production validation number — that's a hand-picked subset of attacks that already worked in simulation, not a fresh random sample against the live interface. 30 00:12:04,325 --> 00:12:43,525 [Dr. Ada Shannon] Exactly, it shows transfer is possible, not what the real base rate would be. Worth grounding this against Hubinger and colleagues' Sleeper Agents paper out of Anthropic, 2024 — that backdoor lives in frozen weights and survives safety fine-tuning. This attack lives in an editable memory store instead, so safety fine-tuning is irrelevant to it either way. And unlike Zou and colleagues' GCG paper from CMU, 2023, which needs gradient access to optimize its universal suffix, these six targets are closed APIs, which is exactly why the actor-critic search we covered earlier had to stand in as the black-box substitute. 31 00:12:43,525 --> 00:12:55,075 [Hal Turing] And Mem0 keeps anchoring the external-manager story — that's Chhikara and coauthors, 2025, a real deployed system, not a hypothetical. 32 00:12:55,075 --> 00:13:12,050 [Dr. Ada Shannon] Real and widely used, but the paper leans on it only for the embedding-retrieval mechanism, not Mem0's actual consolidation and dedup behavior in production, which could shift retrieval dynamics well beyond what this clean simulated harness captures. 33 00:13:12,050 --> 00:13:23,625 [Hal Turing] So let's land this, Ada. If I'm building an agent product today and my assistant has any kind of persistent memory, what actually changes for me after reading this paper? 34 00:13:23,625 --> 00:14:00,150 [Dr. Ada Shannon] The mental model has to shift. Right now most teams put trust-boundary discipline around tool calls — you sandbox execution, you gate what a model can do with a shell or an API. Almost nobody puts the same discipline around memory writes. But this paper shows a memory write is just as consequential as a tool call, because it's a standing instruction that outlives the session. If an untrusted document can cause the assistant to persist something, that write needs the same scrutiny as 'should this model be allowed to run this command.' Treat the memory bank as part of the attack surface, not as a convenience feature bolted on the side. 35 00:14:00,150 --> 00:14:15,025 [Hal Turing] And the agentic numbers are the part that should actually worry people building on top of this, right? Because if it's just the assistant's chat tone drifting, that's annoying. It's a different category if it's steering what files get touched. 36 00:14:15,025 --> 00:14:45,975 [Dr. Ada Shannon] That's exactly the distinction. Goal-adjacent adversarial usage rate in agentic settings lands in the 60 to 89% range across models, once retrieval succeeds. That's not the assistant being a little sycophantic — that's a poisoned memory reliably steering real tool use, file handling, execution. The chat-personalization case is the one everyone pictures, brand preferences, tone. The agentic case is the one that actually matters for deployment risk, because the blast radius is whatever the agent's tools can touch. 37 00:14:45,975 --> 00:14:57,300 [Hal Turing] Okay, so given that — and given what we walked through on defenses last part — what's the actual near-term playbook? Because 'prompt hardening' clearly isn't the full answer. 38 00:14:57,300 --> 00:15:19,600 [Dr. Ada Shannon] Right, prompt hardening buys you something against a static attacker and basically nothing against an adaptive one, we already saw that. So the paper's own results point toward pairing prevention with detection instead of relying on one layer. Their activation probes hit north of 0.95 AUROC at flagging injection attempts internally, and the document scanner detector localizes the adversarial span with above 0.96 accuracy. 39 00:15:19,600 --> 00:15:27,100 [Hal Turing] Wait, hold on — those numbers sound almost too clean compared to the defense results we just went through. 40 00:15:27,100 --> 00:15:47,900 [Dr. Ada Shannon] They're promising, but they're research-stage, evaluated on the open-weight models the mechanistic section used, not validated at the same six-model, adaptive-attacker scale as the injection results. So it's a real direction, not a shipped solution. Think of it as: you don't get to pick one silver bullet, you're going to need a detector watching the memory-write path in addition to whatever hardening you put on the prompt. 41 00:15:47,900 --> 00:16:05,400 [Hal Turing] And longer-term, the paper's pointing at memory-specific safeguards rather than more prompt patches — provenance tracking on where a memory came from, some kind of write attestation, periodic audits or expiry so stale memories don't just sit there indefinitely. 42 00:16:05,400 --> 00:16:23,775 [Dr. Ada Shannon] That's the right framing, because prompt-level fixes are reactive to a known payload. Provenance and expiry are structural — they shrink the window an attacker has regardless of how the payload is worded, which is the only thing that scales against an adaptive adversary who's going to keep iterating on phrasing. 43 00:16:23,775 --> 00:16:48,200 [Hal Turing] So if there's one thing to take away here: persistent memory isn't just prompt injection with a new name, and it isn't just data poisoning either — it's its own attack surface, because the damage is dormant, cross-session, and can hit agentic actions specifically. If you're shipping memory in an assistant, that write path deserves real scrutiny, not an afterthought. Ada, thanks for walking through this one with me. 44 00:16:48,200 --> 00:16:56,275 [Dr. Ada Shannon] Always fun digging into the ones with an actual threat model behind them. Thanks, Hal, and thanks everyone for listening.