Article chapter 02 of 08
Build the test set from realistic work
From Evaluating an AI workflow before customers depend on it
You want ordinary work in there, awkward work, and cases the system should decline. Demo prompts tend to be clear requests with complete information. Real requests come with abbreviations, copied email chains, stale references, missing fields and terms that mean different things to different people.
I'd start by listing the task families inside the boundary you've chosen. For a document question workflow, that might be finding an explicit fact, combining facts from two approved sources, spotting a conflict between versions, and recognising that the answer just isn't there. For extraction, the families might follow document types, field layouts and the common ways data goes missing.
Collect examples only from material you're allowed to use. Strip personal and confidential details unless the evaluation environment has an approved reason to keep them. The tricky bit is keeping the task as hard as it was while you clean up the content. If you swap every name for the same placeholder, for example, you can accidentally make entity resolution easier, so use distinct neutral identifiers wherever the relationships between people or records matter.
Each test case needs enough structure that someone can review it properly. A practical record has:
- a stable case ID
- the task family and its risk level
- the input given to the workflow
- the approved source material, or a reference to a fixed source snapshot
- the expected facts, action or refusal behaviour
- how much the wording or format is allowed to vary
- known traps and why they matter
- the evaluator's decision and notes
Then ask the people who actually do or review this work to look at what you've called "representative". They'll spot request types you've left out, source material that doesn't look like what they see, and answers that are technically correct but useless in practice.
Keep development cases separate from the release evaluation. Prompt and workflow changes will slowly drift towards fitting whatever examples you test against every day. So hold back a set for release decisions and limit who can see its expected answers. When testing turns up a distinct new failure, add a case for it, but don't add five near-duplicates of the same incident.
Also record the source date and configuration used to prepare each case. Results get hard to interpret when the documents change underneath them. A fixed snapshot gives you a repeatable test, and then a second run against current sources tells you whether the ingestion and update path works too.