Article chapter 08 of 08
Keep the evaluation attached to what's running
From Evaluating an AI workflow before customers depend on it
Rerun the evaluation set whenever the prompt, model, sources, tools or post-processing change. Tie each released configuration to a test result so you can tell what evidence supported the version that's live now.
Be careful about adding production cases. You want the shape of the task and the failure, without copying confidential content into a general test set. Confirm you have permission, sanitise the material, fix the expected behaviour and label why the case was added. The set will need maintenance as it grows: remove duplicates, update cases when policies go out of date, and keep older source snapshots if you need them to reproduce a past decision.
Monitor the signals that connect back to the evaluation: abstentions, reviewer edits, overrides, failed tool calls, repeated retries, source misses and tasks that land outside the supported boundary. If those move, I'd go and review cases before deciding anything about model quality. The cause could just as easily sit in source ingestion, permissions, how people are using it, or the interface.
Before widening the release, rerun the held-back set and go through the failures by task family. Check that reviewers still have time to do the review the release depends on. Make sure recovery has actually been exercised and isn't only written up somewhere. Then update the release record with the new boundary and the evidence behind it.
If I were reviewing one of these, I'd finish with five questions:
- Does the current set still look like the work reaching the feature now?
- Can reviewers apply the acceptance rules consistently?
- Have critical and combined failures been tested through the full path?
- Is the human review effort low enough for the expected use?
- Can you detect, contain and recover from the failures you've accepted?
If you're unsure about any of them, I'd keep the release inside its current boundary and add the missing case or operational test before expanding it. Before customers depend on an AI workflow, I want a test set that still looks like the work the feature is getting, with the acceptance rules agreed up front and the results tied to the configuration that's live. That's the evidence I'd go back to each time someone wants to widen the release.