Blog
If software produces a score, the rules need a version
A score can look objective while hiding a changing set of weights, thresholds and assumptions. I want the rule set versioned alongside the result so a reviewer can reproduce what happened, explain it to the person affected and distinguish a changed answer from changed evidence.

A score can look objective while hiding a changing set of weights, thresholds and assumptions. I want the rule set versioned alongside the result so a reviewer can reproduce what happened, explain it to the person affected and distinguish a changed answer from changed evidence. This applies to eligibility checks, risk bands, lead ratings, assessments and internal prioritisation. The output may be one number, but it usually depends on a chain of parsing, defaults, transformations and decisions. If any part changes, the same input can produce a different result. A final score alone cannot explain how the software reached it. A scoring rule set should have an immutable version identifier. That version points to the exact weights, thresholds, question mappings, formulas and default behaviour used at the time. I would keep the executable rules in source control or another system with equivalent history and review. A human-readable specification should sit beside them. The specification explains intent and edge cases; the executable form removes ambiguity about what ran. Both need to move through the same approval process.
A version should change whenever a result could change. Editing a label or correcting a spelling mistake does not necessarily require a new scoring version. Changing a threshold from 60 to 65 does. So does altering a weight, changing how missing values are handled or mapping a new answer to an existing category. Do not reuse a version number after rollback. If version 1.4 is withdrawn and its logic is corrected, release 1.5. An immutable history is easier to reason about than a familiar identifier whose contents changed quietly. Keeping the original submission is necessary, but it may not be sufficient. Scoring often converts raw answers into normalised values before applying rules. Suppose a person enters a date, selects several options and leaves one field blank. The application may parse the date in a timezone, map the options to internal codes and substitute a default for the blank. Reproduction depends on those interpreted values as well as the source payload. A result record can include:
- the raw submission reference and its version;
- the normalised input values used by the scorer;
- the scoring-rule version;
- the application or scorer build identifier;
- the calculation time and relevant effective date;
- the final score, band and reason codes.
Reason codes are especially useful. Instead of trying to reconstruct an explanation from a number, the system records which rules contributed to the result. They should be stable identifiers with plain-language descriptions maintained alongside the rule version. Sensitive inputs still need retention and access controls. Reproducibility does not justify copying personal information into every log. Keep identifiers and protected data in the right stores, with links that authorised reviewers can follow. A reviewer should be able to run the historical inputs through the historical rule set without editing production data. That requires more than a database row containing rules_version = 1.4. Package the scoring logic so a version can be loaded independently. Preserve any lookup tables, category mappings and reference data it used. If a score depends on an external value, such as an exchange rate or published classification, snapshot the value or record the dated source. Using "latest" data prevents a historical decision from being reproduced. The replay command should produce a structured explanation: interpreted inputs, rules evaluated, contributions, exclusions and final output. It should also compare that output with the stored result and flag any difference. Build fixtures from approved examples and boundary cases. For a threshold of 60, include values just below, exactly at and just above it. Include missing fields, malformed values and combinations that activate caps or overrides. Run the fixtures whenever scoring code or reference data changes.
A deterministic scorer should return the same answer for the same versioned inputs. If it does not, identify the variable source and record it. Randomness and model-assisted interpretation need their own captured configuration and stronger review. When a score changes, the first question is what changed. There are several possibilities: the person's evidence changed, somebody corrected an input, the rule set changed or the software previously implemented the rule incorrectly. Store these as separate events. An amended submission should create a new input version linked to the original. A policy update should create a new rule version with an effective date. A software correction needs a release record describing which historical results may be affected. This lets the interface say something precise. "Your result changed because answer 7 was corrected" is different from "the same answers were assessed under rules effective from 1 May". That distinction matters to the person receiving the result and to anybody reviewing consistency. Avoid recalculating every old result silently when new rules are deployed. Historical results should retain the version that produced them. If the organisation decides to reassess them, run a controlled migration that creates new result records and preserves the old ones. A spreadsheet edit can alter decisions for many people, so it deserves normal change control.
The change request should state the proposed rule, rationale, effective date, affected population and expected examples. A reviewer can inspect the human-readable difference and the executable difference. Automated tests show boundary behaviour, while a sample replay shows how real distributions may move without exposing private records in the approval document. Before release, confirm that reporting and explanations understand the new version. A dashboard that compares scores across versions without labelling them may suggest movement where only the rules changed. Exports should include the version and reason codes rather than presenting the number as timeless. Access should also be split. People who can propose a rule change should not automatically be able to publish it to production. Emergency changes still need a recorded approver and a follow-up review. Choose a completed result and try to reproduce it in a clean environment. Retrieve the original inputs, load the recorded rule version and generate the explanation. Compare every intermediate value with the stored record. Any missing dependency uncovered by that exercise belongs in the versioned artefact or result metadata. Once one result can be replayed reliably, turn the procedure into an automated check and run it against a small historical sample before each scoring release.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.