Article chapter 05 of 08
Keep the scoring engine small and boring
I'd put the scoring engine in a small module you can test on its own. It validates inputs, applies rules in a documented order and returns a structured result, and it shouldn't care about the conversation history, the current prompt or which model ran the interview.
A typical request has the methodology version, the assessment run ID and a map of confirmed answer values. The response can include:
- validation errors for missing or malformed inputs
- each rule ID and whether it matched
- intermediate dimension values, if the method uses them
- the final level or profile
- reason codes linked to explanation text
- warnings for incomplete or conflicting evidence
- a hash or stable ID for the calculation inputs
Use decimal or integer arithmetic wherever the methodology has exact thresholds. Define rounding in one place and test values just either side of every boundary. You don't want the interface's number formatting deciding a value that another rule then reads.
Be deliberate about missing data, because zero, false, unknown and not applicable are four different things. If you treat a missing answer as zero, you're quietly penalising incomplete evidence. If you just skip it, a percentage can go up because the denominator shrank. The methodology owner has to choose the behaviour, and the engine should show which behaviour applied in the result.
Tie the explanations to reason codes. It's fine for a model to turn structured reasons into readable prose, but keep the codes and facts underneath in the report. Check the generated text against them before release, especially anywhere the wording could suggest certification, cause and effect, or a comparison the rules never made.
Make recalculation idempotent. Running the same version against the same confirmed answers should give you the same result ID, or an equivalent result recorded as a duplicate of the first. Otherwise a retry can create several scores that look independent.