Article chapter 08 of 08
Keep evaluation attached to the released workflow
From Evaluating an AI workflow before customers depend on it
Rerun the evaluation set when the prompt, model, sources, tools or post-processing change. Tie each released configuration to a test result so the team can tell what evidence supported the version now in use.
Add production cases carefully. Capture the task shape and failure without copying confidential content into a general test set. Confirm permission, sanitise the material, fix the expected behaviour and label the reason the case was added. A growing set still needs maintenance. Remove duplicates, update obsolete policies and retain older source snapshots when they are required to reproduce a past decision.
Monitor signals that connect back to the evaluation: abstentions, reviewer edits, overrides, failed tool calls, repeated retries, source misses and tasks routed outside the supported boundary. A change in these signals should trigger case review, not an automatic conclusion about model quality. The cause may sit in source ingestion, permissions, user behaviour or the interface.
Before expanding the release, rerun the held-back set and review the failures by task family. Check that reviewers still have time to perform the required control. Confirm that recovery has been exercised, rather than merely documented. Then update the release record with the new boundary and evidence.
A practical review can finish with five questions:
- Does the current set resemble the work now reaching the feature?
- Can reviewers apply the acceptance rules consistently?
- Have critical and combined failures been tested through the full path?
- Is human review effort low enough for the expected use?
- Can the team detect, contain and recover from the failures it has accepted?
If one answer is uncertain, keep the release inside its current boundary and add the missing case or operational test before expanding it.