# Evaluation (evals)

> A repeatable test set that measures whether a change to an AI system made it better or worse, rather than just different.

Model outputs are not deterministic and "it seems better" does not survive contact with a second opinion. An eval set is a fixed collection of realistic inputs with known-good answers or scoring criteria, run automatically, so that changing a prompt, a retrieval strategy or a model produces a number you can compare against yesterday's.

Skipping this is the difference between an AI feature that improves and one that drifts. Without a baseline there is no way to tell that the prompt edit which fixed one complaint quietly broke three other cases, and no way to justify a model upgrade beyond enthusiasm. A modest eval set built from real tickets beats an elaborate one built from imagination.

## Related terms

- https://dfieldsolutions.com/en/glossary/hallucination.md
- https://dfieldsolutions.com/en/glossary/llm.md
- https://dfieldsolutions.com/en/glossary/rag.md

---

Source: https://dfieldsolutions.com/en/glossary/llm-evaluation
DField Solutions — Dunakeszi, Hungary — dezso@dfieldsolutions.com
Booking: see https://dfieldsolutions.com/en/contact
