30 April 2026 / Applied AI / 8 chapters

Build fixtures for the rules

From Designing an auditable AI capability assessment

Fixtures are small assessment cases with known inputs and expected rule outcomes. They let subject-matter and technical reviewers discuss the method using examples rather than only reading expressions.

Start with one fixture for each intended level or profile. Then add cases around every threshold, prerequisite and exception. Include incomplete evidence, contradictory answers, not-applicable values and inputs that fail validation. If several rules can produce the same label, cover each path.

A useful fixture contains:

  • a short purpose statement;
  • the methodology version;
  • structured answer values;
  • expected rule matches and non-matches;
  • expected intermediate values;
  • expected result and reason codes;
  • any expected warning or validation error.

Keep raw conversational text out of most scoring fixtures. The scoring suite should test the deterministic boundary directly. Maintain a separate set of mapping examples for the conversation layer: a raw answer, expected structured proposal, ambiguity status and required confirmation. Model behaviour may vary, so evaluate acceptable mapping outcomes rather than exact prose.

When a rule changes, run the full fixture suite and inspect every changed result. Updating expected outputs until tests pass defeats the purpose. Each changed fixture needs an explanation from the methodology owner. Some changes are intended, some reveal a wider consequence, and some show that the rule was implemented incorrectly.

Add property checks where they match the method. If increasing a well-defined positive input should never lower a dimension, test that invariant across generated values. Do not assume monotonic behaviour when prerequisites or penalties make it untrue. Properties belong to the documented method, not to generic scoring intuition.

All articles