Article chapter 03 of 08
Measure the review and exception work
From Budgeting and governing an AI workflow after the prototype
Human review can be a big cost, and it's one you usually don't know well at the start. In a prototype, review tends to be done by attentive project staff on a small set of clean cases. In routine use you get repeated decisions, interruptions and cases that need someone with specialist knowledge.
I'd break review into its actual activities instead of putting one time estimate against every case. A reviewer might read the source material, check a proposed field, look into a conflicting result, correct the output, approve an external action and record a reason. Some cases are a quick confirmation. Others go off to a specialist or back to the person who asked.
During controlled operation, collect:
- when a case went into and came out of each review state
- active handling time, where you can measure it without intrusive monitoring
- wait time, kept separate from work time
- reason codes for escalations and rejections
- which role was needed to resolve the case
- whether the model output was accepted, edited or thrown away
- whether the correction points to a source, rule or product problem
Averages can hide a small group of expensive cases, so keep the distribution and look at the slowest cases one by one. They might turn up unsupported file types, ambiguous policy, poor source documents, or a task that probably shouldn't have been automated in the first place.
Treat the review policy itself as a cost input. Sampling a share of low-risk outputs, reviewing every consequential action and escalating uncertain cases all create different workloads. Budget against the policy that's actually been approved. I wouldn't assume review will shrink on its own as the model gets better.
Reviewers also need calibrating: examples, guidance and some way of settling disagreements. When the method changes, they might need to work through new fixtures before going back to live cases. That's operating work too, especially where consistency matters.
Then there's the queue. If the reviewers you need aren't available, cases sit there or someone in another role gets pulled in. I'd only set service targets once you've measured the staffing it takes to meet them. And if approval is where things get stuck, making the model respond faster won't shorten most cases.