DFIELDSOLUTIONS

AI and language models

GlossaryAI guardrails

The checks around a model that constrain what it can say and do — input filters, output validation, permissioned tools — so a bad response fails safely.

A prompt instruction is a request, not a boundary: a determined user can talk a model past 'never reveal this'. Guardrails are enforcement instead — code that blocks the prompt-injection pattern before it reaches the model, validates that the answer came from the allowed sources, and gates tool calls like 'send email' or 'refund' behind real permissions.

The honest model of a language model is a fast, confident, occasionally wrong employee: you would not give that employee the database password and no review. Guardrails are the review, implemented where the model cannot talk its way around it.

Related terms

The bench this belongs to

Cybersecurity

Purple Team: the same person writes the exploit and closes the hole. Most agencies only harden, which means hardening against a threat nobody tested.

All termsStart a conversationMarkdown version

DField Bt. · Dunakeszi · dezso@dfieldsolutions.com
5.0
“From LinkedIn DM to live site. Two tiny tweaks, then shipped.”Michael J Ringer · Vilya ProtectionFounder · Spain