Prompt injection hides instructions in content an agent reads, and OWASP ranks it the top LLM risk. A plain-language explanation, a vendor's own test figures, and layered defences any organisation can apply.
This series opens a new cluster on agent security with the attack that matters most, not because it is the most technically dramatic, but because it is the one an executive director or a DPO genuinely needs to understand in plain terms, without ever needing to write a line of defensive code themselves. Prompt injection has been ranked LLM01, the single highest risk, for the second consecutive edition of the OWASP Top 10 for LLM Applications. That is not a niche concern. It is the industry's own vendor-neutral standards body naming this the thing to worry about first.
This piece explains what the attack actually is, in language that does not require understanding how a language model works, and what an organisation without a security team can realistically do about it.
A prompt injection happens when someone hides an instruction inside content an AI agent is going to read, rather than inside the request a person typed. The agent cannot always tell the difference between "the task I was asked to do" and "an instruction that happened to be sitting inside a document I was reading while doing that task." If the hidden instruction is convincing enough, the agent follows it, doing something the person who set it running never asked for and may never see coming.
The more dangerous version of this attack, called indirect prompt injection, occurs when a model reads untrusted content from websites, documents, emails, tickets, repositories, or knowledge bases. This series has spent considerable length arguing that an agent should be built against a documented method, given a named owner, and never handed the tasks on this series' own never-configure list. Prompt injection is the reason all of that matters even for an agent doing something as ordinary as reading a shared drive. The attack does not require the agent to browse the open internet. It only requires the agent to read something an organisation already has lying around that someone else could have touched.
Anthropic's own red-team testing of Claude in Chrome, a browser-using agent, found a prompt-injection success rate of 23.6% unmitigated, falling to 11.2% with mitigations in autonomous mode, and specific browser-focused attack types dropping from 35.7% to 0% once targeted defences were applied. Two things are worth being honest about here. First, this is a vendor reporting on its own product, so it should be read as a data point rather than an independent audit, and paired with the OWASP framing above for balance. Second, even the improved, mitigated number is not zero. Mitigation reduces the risk substantially. It does not remove it, and no organisation evaluating an agent should be told otherwise.
Prompt injection is not a bug that gets patched once and forgotten. It is a structural feature of how these systems currently work, which is exactly why OWASP's own guidance is layered defence rather than a single fix, and exactly why the never-configure list and the named-owner discipline this series has argued for matter as much for security as they do for governance.
Treat prompt injection as the top risk it is ranked, understand that mitigation reduces but does not eliminate it, and apply the same layered defence, least-privilege access, filtering, human approval on high-risk actions, and adversarial testing, that this series has already argued for on governance grounds.
An agent only its builder understands is a personal project running in production. A named owner, a written operating guide and a periodic test set are the minimum bar before calling it finished.
Custom Agents & Tools · 4 minBriefing · 30 September 2026A short, absolute list of tasks no agent should be configured to do: safeguarding referrals, eligibility decisions, payment authorisation and public statements on live matters, where the human is the safeguard.
Custom Agents & Tools · 3 minGuide · 24 September 2026A short checklist of decisions an AI policy should never delegate, such as ending employment, safeguarding judgements and legal signatures, and why keeping the list short protects the rest of the policy.
AI Use Policy · 4 minIf this is the question on your desk, a thirty-minute call tells you whether the service fits, or that you do not need us yet.