Article chapter 08 of 08
Look at the production state again after launch
After launch, sources keep changing, permissions can drift and review queues can grow. I'd book the first operational review while the release details are still fresh in everyone's head.
Compare the actual task mix with the boundary you evaluated. People will submit requests nobody anticipated, or use an approved feature as a shortcut for a more consequential decision. Where instructions alone are being ignored, add routing or interface constraints. Unsupported tasks should show up in the data so you can either decline them or assess them properly.
Go through corrections and overrides by reason. Repeated source changes point to content maintenance. Frequent schema failures point to runtime validation or prompt issues. Long review times might mean the interface isn't showing reviewers the evidence they need. These are different operating problems, even if every one of them looks like an unsatisfactory answer to the user.
Reconcile a sample of external actions against the target system. Check that your local confirmed state matches the receiving record and that retries haven't created duplicates. Look at the oldest unresolved run and make sure its owner and next action are still right.
Rerun the access and source-removal tests after permission or collection changes. Rerun the release evaluation whenever the model, prompt, tools, retrieval logic or policy changes, and keep the earlier results attached to the earlier configuration instead of overwriting them.
The review should end with actual decisions, such as:
- Keep, narrow or expand the supported task boundary.
- Fix specific controls before changing volume.
- Add distinct failures to the evaluation set.
- Update source or support ownership.
- Pause any action whose state can't be detected and recovered.
That wrong answer from the start is still the test that matters. Production readiness for applied AI mostly comes down to whether, once that answer has gone into the workflow, the team can see where it went, what it changed and how to put it right. Before you automate a consequential step, take one retained run and check whether someone can work out what changed and pick a safe recovery action from the records alone. If they can't do that yet, I'd keep that step under manual control until the records and the recovery path are fixed.