Article chapter 07 of 08
Set a controlled release decision
From Evaluating an AI workflow before customers depend on it
Write release criteria before the final run. This makes it harder to reinterpret a disappointing result after the fact. The criteria need to reflect task risk, critical failures, review effort and the controls available during release.
Avoid relying on one combined percentage. A high overall result can hide complete failure on a small, consequential task family. Set minimums by family where the risks differ. State which critical failures block release regardless of the aggregate result. Include requirements for permissions, logging, recovery and reviewer capacity.
A controlled release decision should name:
- the users and task families included;
- any excluded inputs, actions or source collections;
- the model, prompt and workflow versions approved;
- the evaluation set and result used as evidence;
- unresolved failure modes and their controls;
- required human review and approval points;
- monitoring signals and who will respond;
- the condition for pausing or rolling back.
The first release should make exposure manageable. Limit the task boundary, eligible users or volume where those controls reduce the consequence of an unseen failure. Make the limitation enforceable in the application. A note in the release plan will not stop an unsupported task reaching the model.
Review residual failures with the people responsible for operations, security and the affected business process. The person who will handle a failed action needs to confirm that the log contains enough information and that the correction path is available. Product approval does not test that recovery work.
Record the decision, including reasons for accepted risks. This creates a baseline for later changes. It also prevents a successful demonstration from becoming informal permission to broaden the workflow beyond the evaluated boundary.