Article chapter 01 of 08
Start with the work the feature will perform
From Evaluating an AI workflow before customers depend on it
An evaluation needs tasks that resemble the work reaching the feature after release. Teams can spend weeks improving a score on tidy examples while the intended workflow remains loosely defined, so begin by drawing a boundary around one piece of work.
Write down the trigger, the inputs available at that point, the output expected, the person or system receiving it, and what happens next. If the feature drafts a response, say whether a person edits it before sending. If it classifies a record, identify which queue or rule depends on that classification. If it retrieves information, record whether the answer supports a decision or merely helps someone find a source.
The same model output can carry very different risk in different workflows. A suggested search term is easy to ignore. A category that changes account access has an operational consequence. The evaluation should follow that consequence far enough to reveal who can detect an error and whether they can repair it.
A short workflow description should answer these questions:
- What starts the task?
- Which sources may the feature read?
- What form must the output take?
- Who reviews or consumes the output?
- Which action can follow automatically?
- What state must be retained for later investigation?
Keep the first boundary narrow. "Answer customer questions" hides too many tasks. Account access, billing explanations and product guidance use different sources and tolerate different mistakes. Split them into separate task families, even if one interface eventually handles all three.
Keep the evaluation tied to the configured workflow: its prompt, retrieval, tools, interface and review step. The result should show whether that full setup can perform the named job under the tested conditions, rather than rank a model in isolation.