Article chapter 03 of 08
Decide what counts as acceptable before you run anything
From Evaluating an AI workflow before customers depend on it
I wouldn't grade generative output against one exact reference sentence unless the wording itself matters. It's more useful to define acceptance in terms of required facts, claims it mustn't make, how it uses sources, format, and what it does when it's unsure.
For each task family, write a small rubric. It can mix yes/no checks (is every requested field there?) with a judgement about whether the answer is usable. Tie the subjective labels to things a reviewer can see. "Good answer" gives a reviewer almost nothing to go on. "States the current account status, cites the approved record and doesn't guess at a reason for the status" is something they can actually check.
Work out how you'll treat partial correctness too. If an answer has four correct facts and one made-up one, averaging it into a decent-looking score would be misleading. Some errors wreck the whole result. I'd mark those as critical failures and report them separately from minor formatting or wording problems.
Say what a good response to uncertainty looks like. Depending on the task, the workflow might need to ask a question, return no answer, offer several candidates, or hand the task to a person. Refusing an answerable low-risk case is annoying. A confident guess on a consequential case can create work that's much harder to unwind.
Fields I'd want in the acceptance criteria:
- required facts or fields
- approved sources and citation requirements
- assumptions the output mustn't make
- what counts as a critical error
- acceptable abstention or clarification behaviour
- the required output schema
- the most review or correction you'll put up with before the result stops being worth having
Before the full run, have two reviewers grade a sample independently. Where they disagree, you've usually found a vague criterion. Talk through the reason for each disagreement and fix the rubric, rather than settling it with a majority vote nobody can explain. If trained reviewers can't apply the definition consistently, the score you get out of it won't support a release decision.
Use automated checks for schema validity, required fields, prohibited tokens, source identifiers and deterministic calculations. Use people for meaning, unsupported implications and whether the output actually helps with the task. Report the two separately, otherwise a clean JSON score can hide a bad answer.