30 April 2026 / Applied AI / 8 chapters

Keep the conversation, the evidence and the scoring apart

From Designing an auditable AI capability assessment

I think of the system as three parts that talk to each other through explicit contracts.

The conversation layer asks the questions, asks the follow-ups it's allowed to ask and helps the respondent get their answer into a usable shape. It can change wording and order within limits you've set, but it shouldn't be quietly handing out points.

The evidence layer keeps the respondent's original words alongside the structured answer proposed from them. A response might become a selected option, a number, a date, a list of documents or an explicit "unknown". Keep the source text next to the structured value so a reviewer can see how you got from one to the other.

The scoring layer only takes the structured values, the assessment definition and the rule version. It gives back dimension results, rule outcomes, reasons and any flags. Same inputs plus same version should give the same result every time, which means a model call has no place in this path if you care about reproducing results.

Put a typed boundary between these parts. For each answer, I'd carry the question ID, the assessment run ID, the raw response, the proposed structured value, the mapping status, a confidence or ambiguity flag where it's useful, and who confirmed it. Don't squash the whole interview into a summary and score the summary. Summaries drop qualifications, and sometimes that qualification is exactly what decides which side of a threshold someone lands on.

Splitting things up also helps when something breaks. If the conversation is down, the completed structured answers are still there. The scoring service can reject an incomplete input without losing the interview, and a reviewer can fix one mapped answer and rerun the same rules without replaying the whole conversation.

All articles