DFIELDSOLUTIONS

Playbook

LabPrompt injection in production agents

Prompt injection is not one bug with one fix. It is five attack categories, each with a different defence, which is why teams that add input sanitisation once get a vulnerability report two weeks after launch.

Prompt injection in production agents

Playbook6 sections · 2 min readBy Dezső Mező.mdLLMPrompt injectionOWASPRAG

This is a condensed version of the studio's full playbook. Each category below is a genuinely different attack with a genuinely different mitigation, and treating them as one problem is the mistake that keeps recurring.

01Direct injection

Someone types an instruction into the chat box aimed at your system prompt. The defences are well known: segment system from user content, mark the instruction hierarchy explicitly, and keep rejection patterns. The failure is not knowing this, it is implementing it once and never evaluating it again.

02Indirect injection via documents

The instruction arrives inside a file the user uploaded or a page the agent fetched. The user never typed anything hostile. Anything the model reads that the user did not author has to be treated as untrusted content, not as instructions.

03RAG-index poisoning

The payload is written into the retrieval index itself, so it surfaces for other users and other questions later. This one survives a deploy, which is what makes it worse than the first two.

04Tool-call abuse

The injection does not aim at the answer, it aims at the tools. Authorisation has to live at the tool boundary and be checked per user, not inferred from a conversation the attacker partly wrote.

05Exfiltration via rendered output

Data leaves through what the answer renders: a link, an image source, a markdown reference pointing at an attacker's host. Validating the answer's content is not enough if you then render it without restriction.

06Evaluate it in CI

All five categories belong in an eval harness that runs on every change, for the same reason unit tests do. A defence nobody re-checks is a defence that quietly stops working.

What to take away

  • Five categories, five defences. One sanitiser covers one of them.
  • Treat everything the model reads but the user did not write as untrusted.
  • Put authorisation at the tool boundary, per user, not in the prompt.
  • Restrict what the answer is allowed to render, not just what it says.
  • Run all five as evals in CI, or the defences rot silently.
We build this for clientsCybersecurity

Read the full write-up

More from the lab

Browse all entries

Want this looked at on your own system?Start a conversation

DField Bt. · Dunakeszi · dezso@dfieldsolutions.com
5.0
“From LinkedIn DM to live site. Two tiny tweaks, then shipped.”Michael J Ringer · Vilya ProtectionFounder · Spain