17 October 2024 / Applied AI / 8 chapters

Define acceptable output before running the model

From Evaluating an AI workflow before customers depend on it

Avoid grading generative output against one exact reference sentence unless the wording itself matters. Two answers can use different words and still support the same action. Define acceptance through required facts, prohibited claims, source use, format and behaviour under uncertainty.

For each task family, write a small rubric. It can include binary checks such as whether every requested field is present, plus a judgement about whether the answer is usable. Keep subjective labels anchored to observable features. "Good answer" gives reviewers little guidance. "States the current account status, cites the approved record and does not infer a reason for the status" is reviewable.

Decide how partial correctness will be treated. An answer containing four correct facts and one fabricated fact should rarely receive an average score that looks acceptable. Some errors invalidate the whole result. Mark these as critical failures and report them separately from minor formatting or wording problems.

Specify an acceptable response to uncertainty. The workflow may need to ask a question, return no answer, present several candidates, or send the task to a person. A refusal on an answerable low-risk case may be inconvenient. A confident guess on a consequential case can create work that is much harder to unwind.

Useful acceptance fields include:

  • required facts or fields;
  • approved sources and citation requirements;
  • disallowed assumptions;
  • critical error conditions;
  • acceptable abstention or clarification behaviour;
  • required output schema;
  • maximum review or correction allowed before the result loses its value.

Have two reviewers independently grade a sample before the full run. Differences expose vague criteria. Discuss the reason for each disagreement and repair the rubric rather than forcing agreement through an unexplained majority vote. If trained reviewers cannot apply the definition consistently, the resulting score will not support a release decision.

Use automated checks for schema validity, required fields, prohibited tokens, source identifiers and deterministic calculations. Use human review for meaning, unsupported implications and whether the output helps with the actual task. Report them separately so a clean JSON score does not hide a bad answer.

All articles