DFIELDSOLUTIONS

Build notes

LabPrompts are code: eval them like code

A prompt that cannot be regression-tested is a script nobody dares to refactor. Versioned, evaluated, diffed — the same hygiene as any deployed artefact.

Prompts are code: eval them like code

Build notes3 sections · 2 min readBy Dezső Mező · Published October 3, 2026.mdLLMEvalsCIPrompting

Shipping an agent means changing the prompt every week. Without an eval harness, each change is a coin flip — the fix for one failure quietly breaks two behaviours nobody checked.

01The eval set is the contract

Fifty to two hundred representative inputs with expected outputs — or at least expected properties — turn 'seems better' into a number. Build it from real transcripts, not from what the team imagines users ask.

02Judge with a judge, spot-check with eyes

An LLM-as-judge scores each run against rubric criteria and scales to every change; a weekly human read of ten random outputs catches what the rubric forgot to ask. Neither alone is enough.

03Version the prompt like the code

Prompt text in version control, eval scores in CI on every change, and a rollback that takes minutes rather than a meeting. The prompt is a deployed artefact and deserves the same hygiene.

What to take away

  • Build the eval set from real transcripts, not guesses.
  • LLM-judge at scale, human spot-check weekly.
  • Prompt changes go through CI like any other code.
  • Keep rollback trivial — prompts deploy too.
We build this for clientsAI automation

More from the lab

Browse all entries

Want this looked at on your own system?Start a conversation

DField Bt. · Dunakeszi · dezso@dfieldsolutions.com
5.0
“From LinkedIn DM to live site. Two tiny tweaks, then shipped.”Michael J Ringer · Vilya ProtectionFounder · Spain