Article chapter 04 of 08
Test the configured workflow as a system
From Evaluating an AI workflow before customers depend on it
A model tested in isolation leaves most of the customer path untested. The deployed feature may add document retrieval, prompt assembly, tool calls, output parsing, business rules and a user interface. Any of those layers can change the result.
Run cases through the same path intended for release. Use the planned model version and parameters. Load the same source index. Apply the real permissions. Exercise the parser and post-processing code. If the interface truncates source excerpts or transforms the response, include that behaviour in the review.
Capture a run record that makes results comparable:
- evaluation set version;
- application and prompt version;
- model identifier and settings;
- source snapshot or index version;
- tool configuration;
- case identifier;
- raw model response;
- parsed or displayed output;
- tool calls and errors;
- elapsed time and review result.
Repeat cases where output can vary. A single successful run may hide instability, while one poor run may exaggerate it. The number of repetitions should follow the consequence and variability of the task. Record every attempt. Keeping only the best result turns evaluation into selection.
Changes should be tested against the whole relevant set. A prompt that improves missing-information cases can make ordinary answers overly cautious. A retrieval change that finds more documents can introduce stale or unauthorised material. Compare case-level changes, not just the aggregate score, and inspect any critical failure even when the total moves in the preferred direction.
Test the surrounding deterministic rules directly as well. Permissions, state transitions, duplicate handling and calculations should have conventional software tests. Generative evaluation does not replace those checks. It adds evidence for the parts whose acceptable output cannot be captured by a simple expected value.