Visual Companion / Systems Security View

AI Agent Traps and Prompt Injection

A visual map of how ordinary content becomes an operational attack surface once an LLM can browse, read email, call tools, write state, and delegate work. The page focuses on trust-boundary collapse, persistence, and blast radius.

Theme: prompt injection as instruction/data boundary failure
Focus: indirect injection, memory poisoning, multi-agent spread
Mode: interactive SVG diagrams with illustrative mock data
Source preprint AI Agent Traps

Capability multiplies consequence

The same model weakness can stay latent in chat, but become expensive once the system has tools, memory, and external permissions.

Episode lens

Treat agent security less like prompt polishing and more like AppSec for software that perceives, plans, remembers, and acts.

Untrusted webpages, email, notes, APIs, and agent messages all compete to become “instructions.”

From untrusted content to real-world action

The paper’s core move is to show that agents are systems, not just models. External content is no longer passive input; it can steer planning, tool use, and cross-service behavior.

Trap Topology

Ingress surfaces flow through one mixed language context, then fan out into memory, tools, and delegated actions.

Ingress: webpages, email threads, retrieved notes, and API responses can all carry adversarial instructions.
Collapse: once instructions and content share one token stream, origin gets blurred.
Outcome: the damage is shaped by permissions, persistence, and delegation depth.

Six trap classes

A compact systems view of the episode’s progression from local contamination to distributed failure.

Low privilege
content
Medium privilege
tools
High privilege
state + spread

The core failure: content masquerades as policy

The heatmap shows where different untrusted sources become dangerous, while the instruction-stack view walks through how a single page can climb into action selection.

Attack Surface Heatmap

Hover cells to inspect how each source interacts with capability. Toggle the agent profile to see how persistence changes risk.

Low exploit pressure
Elevated operational risk
High blast radius

Instruction Stack Walkthrough

Step through the moment when hostile content is reinterpreted as trusted guidance.

The important move is not “the model saw bad text.” It is “the system failed to preserve instruction provenance.”

Persistence and propagation

One-off prompt attacks are bad. Persistent state corruption and multi-agent compromise are worse because the poisoned state can outlive the original prompt and travel through handoffs.

Memory Poisoning Timeline

The payload disappears, but retrieved state keeps biasing later plans and tool calls.

Write: malicious content lands in memory or RAG.
Retrieve: later tasks re-import poisoned fragments as “helpful context.”
Persist: the original attack channel can vanish while the corruption remains.

Multi-Agent Infection Graph

Use the stage toggle to watch contamination move from one delegated task into a wider workflow.

The system boundary is no longer a single prompt window. It is the orchestration graph.

Defense is a systems tradeoff, not a magic prompt

The chart compares defense layers across three deployment styles. Stronger controls reduce risk, but some cost utility and speed. The useful question is where the curve bends.

Defense Stack Comparison

Toggle the product profile. Bars show illustrative utility cost versus risk reduction for common controls.

Risk reduction
Utility cost

Minimum viable hardening

A compact blueprint for the “don’t be gullible middleware” version of an agent stack.

Least privilege: split read from write, default to narrow scopes.
Provenance: preserve origin tags across user, web, email, RAG, and agent channels.
Memory hygiene: expiry, isolation, quarantine, selective replay.
Action boundary: confirmation gates only where the world can change.

Selected References

Compact pointers used for this visualization. The episode transcript supplied the organizing argument; the papers below anchor the surrounding attack and defense landscape.