28 May 2026 / Applied AI

A score needs a methodology readers can inspect

If software gives an organisation a capability score, the reader needs to see what was measured and which rule version produced it. Build that explanation beside the scoring logic so it changes with the method.

Hands using a digital caliper to measure a machined metal part.
Photo: Hans Westbeek (opens in a new tab)

If software gives an organisation a capability score, the reader needs to see what was measured and which rule version produced it. Build that explanation beside the scoring logic so it changes with the method. A number on a dashboard carries more confidence than it deserves. The reader may see 68 out of 100 and assume it came from a stable measurement, even when the result depends on questionnaire wording, missing answers, category weights and thresholds chosen for a particular purpose. Hiding those choices in code makes the score difficult to explain and easy to misuse. The first methodology note should describe the decision the score supports. A capability assessment used to plan internal work has different requirements from a rating used to compare organisations. A score intended to start a workshop may tolerate broad categories. A score tied to funding, access or performance management needs much stronger validation and review. The scope should name the assessed unit and the period covered. Is the result about one team, an organisation, a product or a particular workflow? Does it reflect current practice, a completed project or stated intentions? These distinctions affect what evidence belongs in the assessment.

I would also state what the score does not establish, using plain limits rather than legal-looking boilerplate. A self-reported capability score cannot confirm that controls operate in practice unless the assessment collects and checks evidence of operation. If responses come from one person, the result reflects that person's available information. This context belongs near the score. A methodology link hidden in a footer will not correct an interpretation that the main screen actively encourages. A reader should be able to follow the path from an answer to the displayed result. That requires more than publishing a list of questions. The methodology needs to show which answers receive points, whether some questions are conditional and how unanswered items are treated. Weights need an explanation too. Equal weighting is still a choice. If one area contributes twice as much as another, say why it has more influence on the intended decision. The explanation can be brief, but it should exist outside a developer's memory. For a category score, I record:

  • the questions or evidence included in that category;
  • the rule used to convert each response;
  • any weight or cap applied;
  • the denominator used when an item is not applicable;
  • the rounding applied before display.

These details catch ordinary implementation errors. A missing answer can accidentally become zero in one part of the system and disappear from the denominator in another. Rounding each question before aggregation can produce a different result from rounding the final score. Neither behaviour should be left to an accidental code path. Most assessments add labels or recommendations beside the number. A result may be described as "developing", followed by suggested actions. Those words are part of the methodology because they shape how the reader interprets the score. I keep thresholds, labels and recommendation rules in the same versioned definition as the calculation. The report should not contain a separate block of hand-written conditions that can drift away from the score. If the threshold for a category changes, the explanation and recommendation should change in the same release. Generated narrative needs tighter boundaries. A language model can make a report easier to read, but it should receive the calculated result and approved evidence rather than decide the score itself. The report should distinguish recorded answers, deterministic calculation and generated wording in its internal audit record. Any factual statement in the narrative needs a path back to the assessment evidence. A reviewer should be able to reproduce the number without relying on the generated prose. If the prose contradicts the rules, the product needs a check that catches the mismatch before the report reaches the reader.

A methodology changes over time. Questions are clarified, weights move, thresholds are adjusted and new categories are introduced. Replacing the rules in place makes old reports hard to explain because the same answers may produce a different result today. Each completed assessment should retain the rule version, input answers, relevant evidence, calculated components and displayed result. The version identifier should appear in the report or its accessible details. A date alone is weak because several changes can occur on the same day, and two environments may deploy at different times. When a method changes, decide how existing results behave. They may remain fixed under the original version, be recalculated and clearly marked, or allow the reader to compare both. The choice depends on the purpose of the product. Silent recalculation is risky because a score can move even though the organisation supplied no new evidence. Change notes should describe the practical effect. "Updated scoring logic" tells a reviewer very little. "Questions 4 and 5 now share a category cap, reducing double counting when both describe the same control" gives them something they can check. A scoring engine can be deterministic and still be wrong. I use fixtures that state the inputs, rule version, expected component values, final score, label and recommendations. They run whenever the rule definition or calculation code changes.

Boundary cases deserve their own fixtures. Test values immediately below and above every threshold. Include missing answers, not-applicable items, maximum and minimum results, conditional questions and rounding edges. If categories depend on one another, include the combinations that activate those rules. Someone familiar with the assessment subject should review fixtures as examples of the intended method. Developers can confirm that code matches a rule while missing that the rule itself produces an implausible outcome. Record that judgement separately from the automated test so both forms of review remain visible. Most readers do not need source code or a long technical specification. They need a concise explanation beside the result: purpose, scope, evidence basis, calculation approach, version and important limits. A reviewer may need the question-level rules, weights, thresholds, fixtures and change history. Both views should come from the maintained method rather than two manually synchronised documents. Before releasing a score, take one completed assessment and follow it from source answer to recommendation using only the information the product retains. Any step that requires an undocumented explanation is a gap to close before the methodology is published.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs