Case study · 2026 · AI · Security · Custom software
An AI agent goes to move real money.Built an MCP intent-guard for AI agents. Before the agent transfers money, the system checks intent. Like a colleague asking 'are you sure?'.
MCP Security sits in front of any MCP-compatible agent stack and validates every tool call at the intent level: a secondary classifier model reads the call's purpose, matches it against a YAML policy, allows / blocks, and logs the decision for audit. The studio shipped the intent classifier, the policy engine, and the alerting + audit-log layer.
- Python
- FastAPI
- OpenAI
- Anthropic
- LangChain

Inside the build



Overview
- 1 layer
- Protects every tool
- <1 hr
- Drop-in time
- 100%
- Audit-log coverage
- PI-safe
- Stops prompt-injected calls
Before an AI agent calls a tool (send_email, execute_sql, transfer_funds), MCP Security intercepts. A secondary AI model classifies the intent of the call, matches it against a policy, allows or blocks · logs everything for audit. Drop-in in front of any MCP-compatible agent stack.
What shipped
What it does
- Intent analyser · separate model decides what the call is trying to do
- Policy engine · YAML-based allow/deny rules
- Full audit log · every call, decision, rationale preserved
- Real-time alerts · Slack, PagerDuty, email on suspicious patterns
- Drop-in MCP-compatible · OpenAI Assistants API, Anthropic, LangChain
The problem
- An agent with tools executes whatever the prompt says · prompt injection becomes expensive fast
- Writing custom auth per tool is engineering-days per tool
- No central log of what the AI tried · audit isn't reproducible
- If one tool gets compromised, the others inherit the blast radius
Why it matters
- One layer that protects every tool · no per-tool auth logic needed
- Audit trail ready · fits MNB / NAIH / EU AI Act reporting
- Safe under prompt injection · the intent layer stops malicious calls
- Drop-in · 1 hour to slot into an existing MCP stack
How it shipped
- 01 · BRIEF
What does 'safe' mean for an agent calling tools?
Workshop with the customer's security + AI team to draft the threat model: prompt injection, lateral compromise across tools, exfiltration via legitimate-looking calls. YAML policy DSL fell out of that conversation.
- 02 · BUILD
Intent classifier + policy engine + audit log.
FastAPI middleware fronts the MCP server, calls the secondary model to classify intent, evaluates the YAML policy, returns allow / deny / quarantine. Every decision logged with rationale · replayable for audit.
- 03 · SHIP
Live on customer agent stacks · alerts wired.
Drop-in slot took under an hour per stack · Slack + PagerDuty alerts on suspicious-pattern matches, weekly digest emailed to the security lead.
Stack
Intent
Secondary classifier model
Classifies what each tool call is trying to do · the policy engine never sees raw prompts, only the inferred intent.
Policy
YAML allow / deny / quarantine
Versioned, code-reviewable policy file · changes go through the same PR review as application code.
Audit
Replayable decision log
Every call + decision + rationale stored · auditor can replay any agent run end-to-end.
Alerts
Slack + PagerDuty + weekly digest
Real-time alert on the suspicious patterns plus a weekly digest of trends · security lead reads one email per week.
Case study
“Our AI agents were already touching live systems, and we had no good answer for the auditor. The studio added a layer that checks every action before it happens, like a colleague asking 'are you sure?' before the button is pressed. The audit log and the rule file passed the regulator's review on the first try.”

