Trap Topology
Ingress surfaces flow through one mixed language context, then fan out into memory, tools, and delegated actions.
A visual map of how ordinary content becomes an operational attack surface once an LLM can browse, read email, call tools, write state, and delegate work. The page focuses on trust-boundary collapse, persistence, and blast radius.
Source preprint AI Agent TrapsThe same model weakness can stay latent in chat, but become expensive once the system has tools, memory, and external permissions.
Treat agent security less like prompt polishing and more like AppSec for software that perceives, plans, remembers, and acts.
The paper’s core move is to show that agents are systems, not just models. External content is no longer passive input; it can steer planning, tool use, and cross-service behavior.
Ingress surfaces flow through one mixed language context, then fan out into memory, tools, and delegated actions.
A compact systems view of the episode’s progression from local contamination to distributed failure.
The heatmap shows where different untrusted sources become dangerous, while the instruction-stack view walks through how a single page can climb into action selection.
Hover cells to inspect how each source interacts with capability. Toggle the agent profile to see how persistence changes risk.
Step through the moment when hostile content is reinterpreted as trusted guidance.
The important move is not “the model saw bad text.” It is “the system failed to preserve instruction provenance.”
One-off prompt attacks are bad. Persistent state corruption and multi-agent compromise are worse because the poisoned state can outlive the original prompt and travel through handoffs.
The payload disappears, but retrieved state keeps biasing later plans and tool calls.
Use the stage toggle to watch contamination move from one delegated task into a wider workflow.
The system boundary is no longer a single prompt window. It is the orchestration graph.
The chart compares defense layers across three deployment styles. Stronger controls reduce risk, but some cost utility and speed. The useful question is where the curve bends.
Toggle the product profile. Bars show illustrative utility cost versus risk reduction for common controls.
A compact blueprint for the “don’t be gullible middleware” version of an agent stack.
Compact pointers used for this visualization. The episode transcript supplied the organizing argument; the papers below anchor the surrounding attack and defense landscape.