
Build notes3 sections · 2 min readBy Dezső Mező · Published October 3, 2026.mdLLMEvalsCIPrompting
Shipping an agent means changing the prompt every week. Without an eval harness, each change is a coin flip — the fix for one failure quietly breaks two behaviours nobody checked.
01The eval set is the contract
Fifty to two hundred representative inputs with expected outputs — or at least expected properties — turn 'seems better' into a number. Build it from real transcripts, not from what the team imagines users ask.
02Judge with a judge, spot-check with eyes
An LLM-as-judge scores each run against rubric criteria and scales to every change; a weekly human read of ten random outputs catches what the rubric forgot to ask. Neither alone is enough.
03Version the prompt like the code
Prompt text in version control, eval scores in CI on every change, and a rollback that takes minutes rather than a meeting. The prompt is a deployed artefact and deserves the same hygiene.
What to take away
- Build the eval set from real transcripts, not guesses.
- LLM-judge at scale, human spot-check weekly.
- Prompt changes go through CI like any other code.
- Keep rollback trivial — prompts deploy too.
More from the lab
Browse all entries
RAG that cannot quote a wrong price
Self-hosted n8n or Zapier — where the bill actually differs
MCP servers in production, not in the demo
The invoice pipeline that ended the Friday admin block
Want this looked at on your own system?Start a conversation
