Article chapter 01 of 08
Start with the job the feature will actually do
From Evaluating an AI workflow before customers depend on it
If you're about to put an AI feature in front of customers, the first thing I'd want is a set of test tasks that look like the work it'll get after release. It's easy to spend weeks pushing a score up on tidy examples while nobody has pinned down what the workflow is meant to do, so I'd start by drawing a line around one piece of work.
Write down what triggers the task, what inputs are available at that point, what output you expect, who or what receives it, and what happens next. If the feature drafts a response, does a person edit it before it goes out? If it classifies a record, which queue or rule depends on that classification? If it retrieves information, is the answer feeding a decision or just helping someone find a source?
This matters because the same model output can carry very different risk depending on where it lands. A suggested search term is easy to ignore. A category that changes someone's account access has an operational consequence. I'd follow that consequence far enough to see who could spot an error and whether they'd be able to fix it.
A short workflow description should answer these:
- What starts the task?
- Which sources is the feature allowed to read?
- What form does the output need to take?
- Who reviews or uses the output?
- Which actions can follow automatically?
- What state do you need to keep for later investigation?
Keep that first boundary narrow. "Answer customer questions" hides far too many tasks. Account access, billing explanations and product guidance use different sources and put up with different mistakes, so I'd split them into separate task families even if one interface ends up handling all three.
And keep the evaluation tied to the workflow as you've actually configured it: the prompt, retrieval, tools, interface and review step. You're finding out whether that whole setup can do the named job under the conditions you tested, which a model ranking on its own won't tell you.