30 April 2026 / Applied AI / 8 chapters

Define the assessment contract first

From Designing an auditable AI capability assessment

Before choosing a model or writing prompts, define what the assessment is allowed to conclude. Start with its users, purpose, scope and consequences. A self-guided diagnostic that suggests areas for discussion has a different evidence burden from an assessment used to approve funding, rank organisations or make decisions about people.

Write a short contract covering:

  • the capability or condition being assessed;
  • the population for which the method is intended;
  • the unit being scored, such as an organisation, team, process or response;
  • the evidence the method accepts;
  • how missing, contradictory and uncertain evidence is handled;
  • who may review or amend an answer;
  • the decisions the result may inform;
  • the uses the result does not support.

The contract should also state whether the score is descriptive, comparative or predictive. Do not let report copy imply a stronger meaning than the method provides. A level called "advanced" may sound like an independently validated benchmark when it is only the output of an internal rubric. Define labels in terms of observable criteria and keep unsupported claims out of the interpretation.

Choose the score structure with the methodology owner. Some assessments need weighted dimensions, threshold rules or prerequisite gates. Others may be better represented as a profile with no overall total. Software should implement the approved structure rather than create mathematical precision for its own sake.

Every question should have a reason to exist. Link it to the rule or evidence requirement it supports. If a question cannot affect the result, explanation or review, remove it or mark it as contextual information. This keeps the conversation shorter and makes later changes easier to assess.

All articles