Conventional software is tested by asserting exact outputs. AI systems produce different valid phrasings of the same correct answer, so that approach collapses immediately — and a lot of teams respond by not testing at all. They tweak a prompt, eyeball a few outputs, and ship on vibes.
Evals are the discipline that replaces the vibes.
What an eval set actually is
A collection of representative inputs paired with what a correct response looks like. Run the system against all of them, score the results, and you have a number. Change something, run again, compare numbers.
That is the whole mechanism, and it changes the character of the project: every decision becomes measurable instead of arguable.
Building one that is worth having
Draw from production, randomly. Curated examples flatter the system. Sample real inputs, including the ones that arrived malformed.
Include the hard cases deliberately. Roughly: 60% typical cases, 25% known-difficult, 15% edge cases and adversarial inputs. The easy cases confirm you have not regressed; the hard ones are where improvement happens.
Start smaller than you think. 100–200 well-chosen cases beat 2,000 careless ones, and you can build 150 in a couple of days. Grow it as failures teach you what is missing.
Freeze it. The eval set must not be tuned against casually, or it stops measuring anything. Keep a separate held-out set you touch rarely.
Every production incident should end by adding that case to the eval set. Over a year this is what compounds into a genuinely strong test suite.
What to measure
Different tasks need different scoring, and using the wrong one produces confident nonsense.
| Task | Score with |
|---|---|
| Classification / routing | Accuracy, precision, recall, confusion matrix |
| Extraction | Field-level exact match; partial credit for near-misses |
| Retrieval | Recall@k — was the right passage retrieved at all |
| Generation | Rubric scoring, human or model-judged |
| Agents | Task completion rate, steps taken, cost per run |
For RAG systems, measure retrieval and generation separately. If the correct passage was never retrieved, the generation score tells you nothing useful and you will spend weeks tuning the wrong stage.
LLM-as-judge, used carefully
For open-ended generation, human scoring does not scale, so a model scores the output against a rubric. This works, with caveats that are routinely ignored:
- Validate the judge. Have humans score 50 cases, have the model score the same 50, and measure agreement. If they disagree meaningfully, fix the rubric before trusting any number it produces.
- Use a specific rubric, not "rate this 1–10." Score named dimensions — factual accuracy against context, completeness, format compliance — each with explicit criteria.
- Judges have biases. Models tend to prefer longer answers and their own outputs. Keep a human spot-check running.
Regression testing in CI
Once the eval set exists, wire it into the pipeline. Every prompt change, model version bump, or retrieval tweak runs the suite; a score drop past a threshold fails the build.
This is what makes model upgrades safe. When a new model version ships, you get an answer in twenty minutes instead of a fortnight of anxious manual spot-checks.
Our proposed policy makes this a deployment gate rather than a nice-to-have: GCAS-02 requires evaluating the model together with its prompts, retrieval, memory, connectors, and permissions before it receives authority.
Frequently asked questions
How many test cases do I need?
Enough that a meaningful change moves the number outside normal noise. In practice 100–200 is a workable start for a focused task; below about 50, run-to-run variance swamps real differences. Scale up as you discover failure modes worth encoding.
Who should write the eval set?
Whoever knows what a correct answer looks like — the domain expert, not the engineer. The most valuable week in many projects is a subject-matter expert sitting down and labelling 200 real cases. Engineers tend to unconsciously encode the system's existing behaviour as correct.
How is this different from traditional software testing?
Traditional tests assert exact equality and pass or fail. Evals score a distribution of acceptable outputs and produce a percentage. You are asking "is this good enough, often enough" rather than "is this exactly right." The engineering practices around them — version control, CI, regression gates — are otherwise the same.
Can I use a public benchmark instead?
Public benchmarks tell you about general capability, not about your task. A model topping a leaderboard may perform poorly on your specific documents and taxonomy. Use benchmarks to shortlist candidate models; use your own eval set to choose between them.
How often should evals run?
On every change to prompts, models, retrieval, or chunking — automatically, in CI. Plus a scheduled run against live production samples to catch drift, since your inputs change even when your code does not.
