Model outputs are not deterministic and "it seems better" does not survive contact with a second opinion. An eval set is a fixed collection of realistic inputs with known-good answers or scoring criteria, run automatically, so that changing a prompt, a retrieval strategy or a model produces a number you can compare against yesterday's.
Skipping this is the difference between an AI feature that improves and one that drifts. Without a baseline there is no way to tell that the prompt edit which fixed one complaint quietly broke three other cases, and no way to justify a model upgrade beyond enthusiasm. A modest eval set built from real tickets beats an elaborate one built from imagination.
Related terms
Hallucination
A confident, fluent, entirely invented answer — a citation, a figure or an API that does not exist.
Large language model (LLM)
A model trained on enormous amounts of text to predict what comes next, which turns out to be enough to write, summarise, translate and reason through many tasks.
RAG (retrieval-augmented generation)
Looking up relevant documents first and putting them in the prompt, so the model answers from your material instead of from memory.
The bench this belongs to
AI automationThe repetitive half of your week, handed to software that does not get bored. Inbox triage, follow-ups, reporting, data entry between tools that were never meant to talk.
