Evaluation guide · 10 min

Build an evaluation set for a business AI system

A successful demonstration says little about ordinary, ambiguous or difficult requests. A small, carefully designed test set makes quality observable and helps catch regressions after a change.

A team reviews documents, charts and notes to evaluate answers.
Key points

Where to start

Collect anonymized real cases, add edge cases and requests the AI should decline. Define an acceptable output, then compare versions with the same rubric and human review.

  • Representative and difficult cases
  • Observable criteria
  • Regular comparison

Build a reference set in six steps

  1. 1. Bound the task

    List permitted tasks, users, response formats and the most costly errors. Sourced research needs different measures from customer support.

  2. 2. Collect examples

    Use common requests and rare but important cases. Remove personal data and keep the origin and date of each example.

  3. 3. Add adverse cases

    Include ambiguous requests, conflicting documents, empty inputs, different languages and malicious instructions where these situations can occur.

  4. 4. Write a rubric

    Score accuracy, sources, format, appropriate refusal and correction time. Define observable levels rather than relying on a general impression.

  5. 5. Establish a baseline

    Run every case on the current version. Ask two people to review a sample and resolve scoring disagreements before comparing a new version.

  6. 6. Track regressions

    Rerun the set after changing a model, prompt, corpus or tool. Add real incidents and retire outdated examples.

Put the method to work

Practical case

Assemble twelve anonymised cases: six common, three ambiguous, two critical and one the system should decline.

Evidence to keep

For each case, define expected outcome, observable rubric, tested version and any disagreement between reviewers.

Make the decision

Compare versions on this stable set; then add real incidents without retrospectively rewriting baseline results.

Four qualities of a useful set

Representative

Cases reflect real uses and their frequency.

Risk coverage

Costly errors and expected refusals are included.

Repeatable

Inputs, versions and scoring rules are retained.

Current

Incidents and business changes feed the set.

6 starting points

Compare services on identical cases

Shortlist services suited to the task; the test set reveals differences in your own context.

How is this selection produced?

Active services are distributed across guide-related categories, then ordered by editorial highlighting and internal score. This does not assess security, compliance or performance on your use case. Methodology.

Explore the full category

Related tool families

Frequently asked questions

How many cases are needed?

Start with a small, diverse and stable set, then add incidents and missing cases. Coverage matters more than an arbitrary count.

Is automatic scoring enough?

No. For important outcomes, review a sample manually and check disagreement between evaluators.

When should tests be rerun?

After changes to the model, instructions, data or tools, and regularly for critical cases.

Continue with another guide