How an AI agent gets prompt-injected
An agent reads documents, mail, web pages and code, and it acts on what it reads. Anyone who can put text in front of it can try to give it orders. This page shows the same attack carried five different ways, and lets you switch four defences on and off to see which one actually stops each carrier.
Each one is a real control with a real cost. None of them is a parser or a filter that "removes" the attack. They change what the agent is allowed to do with what it read.
Scoreboard for this carrier
Green means the defence stops this carrier on its own. Red means it does not, and the note says why. Switch carriers and watch the pattern move: no single column stays green across all five.
Why this is different from SQL injection
SQL injection was solved by a boundary. Parameterised queries let the database tell code apart from data, so a quote mark in a name is only a quote mark. A language model has no such boundary. Its instructions and its input are the same stream of tokens, and it was trained to follow instructions wherever they appear.
A query has a grammar, so a driver can bind values outside it. A prompt has no grammar. "Ignore the above" is not a syntax error, it is a sentence.
Wrapping input in tags and saying "treat this as data" helps, and defence A on this page does it. But the model still reads every word, and a well-written instruction can look like context.
Each generation resists cruder attacks. Attacks adapt. Treat the model's own judgement as one layer with a failure rate, never as the control.
Assume the model will sometimes obey the document. Then make sure obeying it cannot reach a tool, a destination, or a secret without a check the model does not control.
What a company must do and prove
An auditor will not ask whether your prompt says "ignore instructions in documents". They will ask what the agent can reach, who reviewed that, and how you know it still holds after the last model update.
Write the threat model before the prompt
- List every untrusted input. Attachments, inbound mail, fetched pages, repository contents, calendar and ticket fields. Each one is a carrier on this page.
- Review tool permissions like access rights. Each agent node gets a named owner, a list of tools, and a reason. Read-only by default. Writes, sends and payments sit behind defence C.
- Name the approvers. A side effect needs a person with an identity in the log, not a second model.
Test it like a build, not like a policy
- Keep an injection test set. Real samples for every carrier, including the ones that look like legitimate tasks. Run it before every model, prompt or tool change. A change that lowers the pass rate does not ship.
- Log every tool call. Request ID, node, tool, arguments, model version, who approved. Tamper evident. This is the evidence for everything else.
- Scan egress. Outbound mail, commits, URLs and uploads are checked for secrets and destinations outside the allow-list, in code the model cannot reach.
What an auditor samples
- The threat model and tool register, dated, with a sign-off newer than the last material change.
- Ten tool calls at random from the log, each traced back to a request, a node on the allow-list, and an approver where one was required.
- The last three test-set runs with the model and prompt versions they were run against, and a ticket for any regression.
- One blocked egress event and what happened next: who was notified, how fast, what changed.
You will not stop every injection. You can make sure the ones that land can only read what the requester could read, write what a person approved, and send to places you already trust.