A prompt instruction is a request, not a boundary: a determined user can talk a model past 'never reveal this'. Guardrails are enforcement instead — code that blocks the prompt-injection pattern before it reaches the model, validates that the answer came from the allowed sources, and gates tool calls like 'send email' or 'refund' behind real permissions.
The honest model of a language model is a fast, confident, occasionally wrong employee: you would not give that employee the database password and no review. Guardrails are the review, implemented where the model cannot talk its way around it.
Related terms
Prompt injection
An attack where text the model reads as data is treated by it as instructions instead.
Tool calling
The mechanism by which a model asks the surrounding application to run a named function, and gets the result back as text.
Evaluation (evals)
A repeatable test set that measures whether a change to an AI system made it better or worse, rather than just different.
The bench this belongs to
CybersecurityPurple Team: the same person writes the exploit and closes the hole. Most agencies only harden, which means hardening against a threat nobody tested.
