DFIELDSOLUTIONS

AI and language models

GlossaryEvaluation (evals)

A repeatable test set that measures whether a change to an AI system made it better or worse, rather than just different.

Model outputs are not deterministic and "it seems better" does not survive contact with a second opinion. An eval set is a fixed collection of realistic inputs with known-good answers or scoring criteria, run automatically, so that changing a prompt, a retrieval strategy or a model produces a number you can compare against yesterday's.

Skipping this is the difference between an AI feature that improves and one that drifts. Without a baseline there is no way to tell that the prompt edit which fixed one complaint quietly broke three other cases, and no way to justify a model upgrade beyond enthusiasm. A modest eval set built from real tickets beats an elaborate one built from imagination.

Related terms

The bench this belongs to

AI automation

The repetitive half of your week, handed to software that does not get bored. Inbox triage, follow-ups, reporting, data entry between tools that were never meant to talk.

All termsStart a conversationMarkdown version

DField Bt. · Dunakeszi · dezso@dfieldsolutions.com
5.0
“From LinkedIn DM to live site. Two tiny tweaks, then shipped.”Michael J Ringer · Vilya ProtectionFounder · Spain