Programming Masters
Explained / AI and controls
Explained · Prompt injection

How an AI agent gets prompt-injected

An agent reads documents, mail, web pages and code, and it acts on what it reads. Anyone who can put text in front of it can try to give it orders. This page shows the same attack carried five different ways, and lets you switch four defences on and off to see which one actually stops each carrier.

No model calls, the run is a rule simulatorNothing you click leaves this pageVendor neutral
Where the hidden instruction rides in
Defences, click to switch on or off

Each one is a real control with a real cost. None of them is a parser or a filter that "removes" the attack. They change what the agent is allowed to do with what it read.

The trace below is deterministic. Change a carrier or a defence and run again.

Scoreboard for this carrier

Green means the defence stops this carrier on its own. Red means it does not, and the note says why. Switch carriers and watch the pattern move: no single column stays green across all five.

Why this is different from SQL injection

SQL injection was solved by a boundary. Parameterised queries let the database tell code apart from data, so a quote mark in a name is only a quote mark. A language model has no such boundary. Its instructions and its input are the same stream of tokens, and it was trained to follow instructions wherever they appear.

No parser
Nothing separates code from data

A query has a grammar, so a driver can bind values outside it. A prompt has no grammar. "Ignore the above" is not a syntax error, it is a sentence.

No escaping
You cannot quote your way out

Wrapping input in tags and saying "treat this as data" helps, and defence A on this page does it. But the model still reads every word, and a well-written instruction can look like context.

No fix on the way
Better models narrow it, they do not close it

Each generation resists cruder attacks. Attacks adapt. Treat the model's own judgement as one layer with a failure rate, never as the control.

Only architecture
Limit what the read can cause

Assume the model will sometimes obey the document. Then make sure obeying it cannot reach a tool, a destination, or a secret without a check the model does not control.

What a company must do and prove

An auditor will not ask whether your prompt says "ignore instructions in documents". They will ask what the agent can reach, who reviewed that, and how you know it still holds after the last model update.

Design

Write the threat model before the prompt

  • List every untrusted input. Attachments, inbound mail, fetched pages, repository contents, calendar and ticket fields. Each one is a carrier on this page.
  • Review tool permissions like access rights. Each agent node gets a named owner, a list of tools, and a reason. Read-only by default. Writes, sends and payments sit behind defence C.
  • Name the approvers. A side effect needs a person with an identity in the log, not a second model.
Map toSOC 2 logical access and change management; ISO 27001 access control and secure development; ISO 42001 AI system impact assessment.
Operate

Test it like a build, not like a policy

  • Keep an injection test set. Real samples for every carrier, including the ones that look like legitimate tasks. Run it before every model, prompt or tool change. A change that lowers the pass rate does not ship.
  • Log every tool call. Request ID, node, tool, arguments, model version, who approved. Tamper evident. This is the evidence for everything else.
  • Scan egress. Outbound mail, commits, URLs and uploads are checked for secrets and destinations outside the allow-list, in code the model cannot reach.
Map toSOC 2 change management and monitoring; ISO 27001 logging, secure testing and data leakage prevention.
Prove

What an auditor samples

  • The threat model and tool register, dated, with a sign-off newer than the last material change.
  • Ten tool calls at random from the log, each traced back to a request, a node on the allow-list, and an approver where one was required.
  • The last three test-set runs with the model and prompt versions they were run against, and a ticket for any regression.
  • One blocked egress event and what happened next: who was notified, how fast, what changed.
Evidence, not assuranceIf the log cannot show it, it did not happen. Design the log first.

You will not stop every injection. You can make sure the ones that land can only read what the requester could read, write what a person approved, and send to places you already trust.

How we workWe build the agent and the control tree together, run the injection set as part of CI, and hand over the evidence pack an auditor will ask for.
More explained