Blog
A conversational assessment still needs deterministic scoring
Keep the conversation and the scoring engine separate. The conversation captures evidence; a versioned rule set calculates the result and leaves a reviewer able to reproduce it.

Keep the conversation and the scoring engine separate. The conversation captures evidence; a versioned rule set calculates the result and leaves a reviewer able to reproduce it. A conversational interface can make an assessment easier to complete. It can clarify a question, accept an answer in the person's own words and ask for missing detail. That flexibility becomes a problem if the language model also decides the score without a stable method behind it. The examples here describe a composite architecture. The exact rules will depend on the assessment and require review by the people accountable for its use. Start with the result the assessment must produce and the evidence needed for each part. Define dimensions, allowed responses, weights, thresholds, exclusions and rules for incomplete information before designing the conversation. Each rule should have a stable identifier and a version. If a response maps to an option worth three points, record the mapping and value in data or code that can be tested. Do not bury the rule in a system prompt where a wording change can alter behaviour without appearing as a methodology change. Keep interpretation narrow. Some questions can use explicit choices and need no model classification. Free-text answers may require mapping to a controlled option, extracting a date or identifying whether required evidence was supplied. Specify the allowed output and what happens when the answer does not fit.
Also define when the system must ask a person to review. Ambiguous evidence, conflicting answers and high-consequence outcomes should have an explicit route rather than a lower confidence score hidden inside the total. The conversation layer should know which assessment item is active, what evidence is required and which valid responses the scoring contract accepts. It can phrase the question naturally and acknowledge the answer, but its output to the assessment record should be structured. Store the original response alongside the proposed mapping. If the person selects an explicit option, preserve that selection. If a model maps free text, retain the model output, confidence or reason code, relevant prompt version and the exact controlled option chosen. Do not make a user see implementation language such as "mapped to option B" in the chat. The interface can respond normally and ask a focused follow-up. The structured detail belongs in the review record. A fallback matters. When a response does not provide enough evidence, the conversation should ask a clarifying question or mark the item unresolved. Forcing every answer into the nearest option creates tidy data at the expense of accuracy. Keep personal and sensitive information to what the assessment genuinely requires. The conversational style can encourage people to disclose more than a fixed form. Instructions and storage controls should limit what is collected and explain how it will be used.
Once the evidence record contains controlled responses, pass it to a deterministic scoring function. The same inputs and rule version should always produce the same result. The scoring output should include more than a total. Return the rule version, dimension scores, rules applied, excluded or unresolved items, threshold decisions and any manual overrides. This gives the explanation layer facts it can present without recomputing the method in prose. Use ordinary code for arithmetic and branching. Validate inputs against a schema. Reject unknown option identifiers and rule versions. Handle missing responses deliberately rather than allowing language-model defaults to decide whether they count as zero, neutral or incomplete. If an assessment includes calculations that change by date, region or participant type, make those conditions explicit inputs. A result should not depend on contextual facts that were available in the conversation but never entered the scoring record. Methodologies change. A weight may be corrected, a question may be replaced or a threshold may be adjusted after review. Existing results need to retain the rules that produced them. Publish a new rule version rather than editing the old one in place. Record its effective date, author or approving role, reason for change and migration decision. Some assessments should keep historical results untouched. Others may need recalculation, but that should create a new result linked to the previous one.
Version the question and mapping instructions as well. If two conversation prompts collect materially different evidence for the same rule, the score version alone does not explain the result. A reviewer should be able to select a completed assessment and retrieve its original answers, structured mappings, scoring inputs, exact rule version and calculation output. If a manual change was made, retain the previous value, reason, actor and time. Create a compact set of assessment fixtures early. Include clear low and high cases, boundary values, incomplete answers, contradictory responses, irrelevant free text and language that could plausibly map to more than one option. For each fixture, define the expected structured evidence and deterministic result. Test the conversation mapping separately from scoring. This makes it clear whether a failure came from understanding the answer or applying the method. Run the same scoring fixtures whenever rules change. Add a comparison report that shows which expected results moved and which rule caused the change. A methodology update that changes unrelated dimensions needs investigation before release.
Also test the review interface. Give a reviewer a completed case without access to the chat transcript view and check whether the evidence record and rule trace are enough to reproduce the score. Then inspect the original wording where the mapping is disputed. A result explanation should use the saved dimension scores, applied rules and cited responses. A language model may help turn that structure into readable text, but it should not add reasons that are absent from the record. Keep the methodology available in plain language. State what was measured, which version applied, how incomplete items were handled and whether any response required manual review. Avoid presenting a capability score as a measurement of things the questions did not cover. The practical release check is reproducibility. Take a completed assessment, remove the generated narrative and run its stored scoring inputs through the recorded rule version. The total, dimensions and threshold decisions should match. If they do not, hold the result and fix the record or engine before improving the conversation.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.