Article chapter 01 of 08
Decide what the assessment is allowed to say
Before you pick a model or write a prompt, work out what the assessment is allowed to conclude, who uses it and what happens to people because of the result. A self-guided diagnostic that suggests a few areas to talk about needs much less evidence behind it than something used to approve funding, rank organisations or make decisions about people.
I'd write a short contract that covers:
- the capability or condition you're assessing
- who the method is meant for
- what's being scored (an organisation, a team, a process, a single response)
- what evidence the method accepts
- how you'll treat missing, contradictory and uncertain evidence
- who can review or change an answer
- which decisions the result can feed into
- which uses it doesn't support
It should also say whether the score is descriptive, comparative or predictive, and the report copy shouldn't hint at anything stronger. A level called "advanced" can read like an independently validated benchmark when it's really an internal rubric's output. Define each label by the criteria someone could actually observe, and leave out claims the method can't back up.
Work out the score structure with whoever owns the methodology. Some assessments need weighted dimensions, threshold rules or prerequisite gates. Others are better shown as a profile with no overall total at all. The software implements whatever gets approved, without adding decimal places just because it can.
Every question needs a reason to be there, so link each one to the rule or evidence requirement it feeds. If a question can't change the result, the explanation or the review, either drop it or mark it as context only. That keeps the conversation shorter and makes later changes easier to reason about.