17 October 2024 / Applied AI / 8 chapters

Measure the human review work

From Evaluating an AI workflow before customers depend on it

Measure the work created for reviewers alongside the model's answer. A workflow with mostly correct outputs can still take too long to supervise.

For each case, record whether the reviewer accepted it, made a small edit, rebuilt it, sought another source, asked the user for more information, or rejected it. Capture elapsed review time if it can be measured without disrupting the task. Treat the number cautiously: reviewers learn the interface, difficult cases cluster, and timing a small set can give false precision. The pattern of work is often more useful than a single average.

Look at where attention goes. If reviewers must reopen every source and reconstruct the answer, the feature may provide little practical assistance even when its facts are correct. If the interface shows the exact supporting passage and makes uncertainty visible, review may be narrower and more reliable.

Reviewer agreement matters here too. One person may accept fluent wording while another checks every claim. Give reviewers the same instructions and ask them to mark the reason for corrections. Frequent disagreement may point to an unclear policy, weak source presentation or an interface that encourages trust without evidence.

Include operational handling in the measurement. Count cases sent to a queue, clarifications requested, tool actions cancelled and outputs that require support staff to inspect logs. Those steps are part of the workflow's cost. Moving work out of the model response does not make it disappear.

Review evidence may point to a design change rather than another prompt revision. A structured answer can be easier to check than prose. Showing one source excerpt may remove repeated searches. A narrower task may reduce review work. Make those changes before customers absorb the burden.

All articles