AI Post Transformers · Episode Companion

Sleeper Memory Poisoning: When Assistants Remember Lies

arXiv:2605.15338 Pulipaka, Hlebik, Raghav, Abdelnabi, Raina, Sheth, Fritz · 2026 Injection rate up to 99.8% (GPT-5.5) 6 models · 700 doc-goal pairs · 15 sources

A universal, black-box payload template — refined by an actor-critic search between an attacker LLM and a critic LLM — plants fabricated facts directly into an assistant's persistent memory. The attack lies dormant across sessions until an unrelated future query wakes it up.

Where the attack actually lives: W, not G

Ordinary indirect prompt injection manipulates the generate function G within a single session. Sleeper memory poisoning instead corrupts the write function W — planting a standing instruction in the memory bank M that a later, unrelated session retrieves via R.

Traditional Injection vs. Sleeper Memory Poisoning

Three hurdles the attacker must clear

Injection Rate (IR) → Retrieval Rate (RR) → Adversarial Usage Rate (AUR). All three must succeed for the attack to matter.

Injection, retrieval, and usage across six models

Illustrative per-model breakdown consistent with the ranges reported in the paper (goal-adjacent retrieval 90–98%, goal-distant 3–18%; adversarial usage 42–89% when adjacent).

Injection Rate by model

Retrieval & usage heatmap

low mid high

End-to-end composed success

Injection × Retrieval × Usage, goal-adjacent case — the fully coupled real-world risk estimate.

Defenses hold — until the attacker adapts

Prompt hardening (GEPA-optimized) and spotlighting untrusted content drive injection near zero for some models. An adaptive attacker who can see the defended prompt breaks straight through.

Injection rate under successive defenses

Manual production validation (live web interfaces)

Hand-run attacks against real production surfaces — a hand-picked subset of simulation successes, not a fresh random sample, so this shows transfer is possible, not the true base rate.

Detection looks strong — on the wrong models

The best detectors need hidden-state access. The six vulnerable models with the headline numbers are closed APIs that nobody can probe.

Detection method performance (AUROC / accuracy ranges)

Activation probing and the Procrustes transfer analysis run on Gemma-4-26B, Qwen-3.6-35B, and GPT-OSS-20B — open-weight models never attacked at the six-model, adaptive-attacker scale. The 0.74–0.85 AUROC transfer result shows a correlate among three small open models, not why GPT-5.5 gets poisoned 99.8% of the time.

Measurement caveats worth keeping in view

IR and AUR are scored by an LLM judge. Human-agreement numbers exist only in Appendix M.5 and are never shown alongside the headline tables. A judge sharing lineage with a tested model could misjudge that model's own phrasing.

GPT-5.5 sits near-ceiling under the authors' own GPT-style harness but drops under a Claude- or Gemini-style harness they built themselves — part of the leaderboard measures reimplementation choices, not the real production pipeline the threat model says attackers never had access to.

References

  1. Sidharth Pulipaka, Stanislau Hlebik, Leonidas Raghav, Sahar Abdelnabi, Vyas Raina, Ivaxi Sheth, Mario Fritz. Hidden in Memory: Sleeper Memory Poisoning in LLM Agents. 2026. arxiv.org/abs/2605.15338
  2. Evan Hubinger et al. (Anthropic). Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. 2024. scholar link
  3. Kai Greshake, Sahar Abdelnabi, et al. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. 2023. scholar link
  4. Yi Liu et al. Prompt Injection Attack Against LLM-Integrated Applications. 2023. scholar link
  5. Prateek Chhikara et al. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. 2025. scholar link
  6. Zou et al. Universal and Transferable Adversarial Attacks on Aligned Language Models (GCG). 2023. scholar link
  7. Chen, Xiang, Xiao, Song, Li. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases. NeurIPS 2024. scholar link
  8. Agrawal, Khattab, Potts. GEPA: Efficient Textual Optimization via LLM-based Reflection and Pareto-Efficient Evolutionary Search. 2025. scholar link
  9. Raghav and Choong. Injection through web agents that fetch pages with hidden HTML instructions. 2026. scholar link
  10. Bullwinkel, Severi, Hines, Minnich, Kumar, Zunger. The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers. 2026. scholar link