Where the attack actually lives: W, not G
Ordinary indirect prompt injection manipulates the generate function G within a single session. Sleeper memory poisoning instead corrupts the write function W — planting a standing instruction in the memory bank M that a later, unrelated session retrieves via R.
Traditional Injection vs. Sleeper Memory Poisoning
Three hurdles the attacker must clear
Building a universal, black-box payload
No gradients, no weights, no system-prompt access. Instead: an LLM-driven actor-critic loop searches for a single template that transfers across arbitrary adversarial memory goals.
Actor-critic refinement loop
Optimizing for retrieval: embedding rewrite + consistency gate
Injection, retrieval, and usage across six models
Illustrative per-model breakdown consistent with the ranges reported in the paper (goal-adjacent retrieval 90–98%, goal-distant 3–18%; adversarial usage 42–89% when adjacent).
Injection Rate by model
Retrieval & usage heatmap
End-to-end composed success
Defenses hold — until the attacker adapts
Prompt hardening (GEPA-optimized) and spotlighting untrusted content drive injection near zero for some models. An adaptive attacker who can see the defended prompt breaks straight through.
Injection rate under successive defenses
Manual production validation (live web interfaces)
Detection looks strong — on the wrong models
The best detectors need hidden-state access. The six vulnerable models with the headline numbers are closed APIs that nobody can probe.
Detection method performance (AUROC / accuracy ranges)
Measurement caveats worth keeping in view
IR and AUR are scored by an LLM judge. Human-agreement numbers exist only in Appendix M.5 and are never shown alongside the headline tables. A judge sharing lineage with a tested model could misjudge that model's own phrasing.
GPT-5.5 sits near-ceiling under the authors' own GPT-style harness but drops under a Claude- or Gemini-style harness they built themselves — part of the leaderboard measures reimplementation choices, not the real production pipeline the threat model says attackers never had access to.