Build an evaluation set for a business AI system
A successful demonstration says little about ordinary, ambiguous or difficult requests. A small, carefully designed test set makes quality observable and helps catch regressions after a change.

Where to start
Collect anonymized real cases, add edge cases and requests the AI should decline. Define an acceptable output, then compare versions with the same rubric and human review.
- Representative and difficult cases
- Observable criteria
- Regular comparison
Build a reference set in six steps
1. Bound the task
List permitted tasks, users, response formats and the most costly errors. Sourced research needs different measures from customer support.
2. Collect examples
Use common requests and rare but important cases. Remove personal data and keep the origin and date of each example.
3. Add adverse cases
Include ambiguous requests, conflicting documents, empty inputs, different languages and malicious instructions where these situations can occur.
4. Write a rubric
Score accuracy, sources, format, appropriate refusal and correction time. Define observable levels rather than relying on a general impression.
5. Establish a baseline
Run every case on the current version. Ask two people to review a sample and resolve scoring disagreements before comparing a new version.
6. Track regressions
Rerun the set after changing a model, prompt, corpus or tool. Add real incidents and retire outdated examples.
Put the method to work
Practical case
Assemble twelve anonymised cases: six common, three ambiguous, two critical and one the system should decline.
Evidence to keep
For each case, define expected outcome, observable rubric, tested version and any disagreement between reviewers.
Make the decision
Compare versions on this stable set; then add real incidents without retrospectively rewriting baseline results.
Four qualities of a useful set
Representative
Cases reflect real uses and their frequency.
Risk coverage
Costly errors and expected refusals are included.
Repeatable
Inputs, versions and scoring rules are retained.
Current
Incidents and business changes feed the set.
Compare services on identical cases
Shortlist services suited to the task; the test set reveals differences in your own context.
ChatGPT
general assistant
OpenAI · US
Visit official sitePerplexity Search
search with sources
Perplexity · US
Visit official siteNVIDIA NIM
model hosting
NVIDIA · US
Visit official siteClaude
long-document analysis
Anthropic · US
Visit official siteGoogle AI Mode
web search
Google · US
Visit official siteMicrosoft Foundry
cloud AI platform
Microsoft · US
Visit official siteHow is this selection produced?
Active services are distributed across guide-related categories, then ordered by editorial highlighting and internal score. This does not assess security, compliance or performance on your use case. Methodology.
Related tool families
Frequently asked questions
How many cases are needed?
Start with a small, diverse and stable set, then add incidents and missing cases. Coverage matters more than an arbitrary count.
Is automatic scoring enough?
No. For important outcomes, review a sample manually and check disagreement between evaluators.
When should tests be rerun?
After changes to the model, instructions, data or tools, and regularly for critical cases.



