17 October 2024 / Applied AI / 8 chapters

Build the evaluation set from representative tasks

From Evaluating an AI workflow before customers depend on it

Include ordinary work, awkward work and cases the system should decline. Prepared demonstration prompts tend to over-represent clear requests with complete information. Production requests also contain abbreviations, copied email chains, stale references, missing fields and terms that mean different things to different people.

Start by listing task families inside the chosen boundary. For a document question workflow, those families might include locating an explicit fact, combining facts from two approved sources, identifying a conflict between versions, and recognising that the answer is absent. For extraction, the families may follow document types, field layouts and common forms of missing data.

Collect examples from material that the team is permitted to use. Remove personal information and confidential details unless the evaluation environment has an approved reason to retain them. Preserve the difficulty of the task while sanitising the content. Replacing every name with the same placeholder can accidentally make entity resolution easier, so use distinct neutral identifiers where relationships matter.

Each evaluation record should carry enough structure to support review. A practical record includes:

  • a stable case identifier;
  • the task family and relevant risk level;
  • the input presented to the workflow;
  • approved source material or a reference to a fixed source snapshot;
  • the expected facts, action or refusal behaviour;
  • permitted variation in wording or format;
  • known traps and why they matter;
  • the evaluator's decision and notes.

Ask the people who perform or review the work to inspect what the team has called "representative". They can identify omitted request types, unrealistic source material and examples where the proposed answer would be technically correct but operationally useless.

Separate development cases from the release evaluation. Prompt and workflow changes will gradually fit the examples used every day. Keep a held-back set for release decisions, and limit access to its expected answers. Add new cases when production-like testing uncovers a distinct failure, rather than creating several near-duplicates of the same incident.

Record the source date and configuration used to prepare each case. Evaluation results become hard to interpret when the documents change underneath them. A fixed snapshot gives the team a repeatable test. A second run against current sources can then test whether the ingestion and update path also works.

All articles