Article chapter 08 of 08
Build review and correction into the product
The point of all this reproducibility is giving people a practical way to get a result corrected. Where the process allows it, show the respondent or an authorised reviewer the accepted answer values before final scoring. Use plain labels and keep the original response close by.
For each result, show:
- the methodology version and assessment date
- the dimensions or rules that contributed
- the evidence references used
- anything missing or uncertain
- what the result means and where its limits are
- how to ask for a correction
A correction should say which answer is changing, the old and new values, the reason, who made it and any supporting evidence. Recalculate under the same methodology version unless the correction process specifically allows a version change. Keep both results and mark which one is current.
Before publishing, I'd run four separate reviews. A methodology review covers the questions, the evidence standard, the rules and how results are interpreted. A technical review covers validation, calculation and access control. A content review checks that neither the conversation nor the report claims more than the method supports. A privacy review covers collection, retention, export and deletion.
These are the questions I'd want answered at the release check:
- Can a reviewer reproduce every fixture without calling a model?
- Can the system tell changed evidence apart from changed rules?
- Does every explanation shown trace back to reason codes and confirmed inputs?
- Are missing and not-applicable values handled the way the methodology says?
- Can a respondent fix a mapping without restarting the interview?
- Does recalculation keep the original result and the reason for the change?
- Are raw responses and evidence covered by proper access and retention rules?
- Can support trace a report back to one assessment run without seeing unrelated records?
Then do one last dry run with an ambiguous answer, a corrected mapping, one unknown item and a value sitting right on a boundary. Request the audit export and hand it to an authorised reviewer. If they can reproduce the score from the export and explain each rule that contributed, without another model call, you can start controlled use.
The respondent who got a score and a polished report and couldn't find out why is who all of this is for. Making an AI capability assessment auditable comes down to keeping the AI on the conversation and the evidence, and leaving the score to a small deterministic engine that runs a versioned ruleset over confirmed answers. If you're building one, I'd start with the contract for what the assessment is allowed to say and a first set of fixtures, then check that every score in the report can be traced to a rule version, a reason code and the answers it came from.