Article chapter 05 of 08
Go looking for the failure modes
From Evaluating an AI workflow before customers depend on it
If you wait for the test set to stumble onto failures by chance, you'll have predictable gaps. Use the workflow description from the start to work out where inputs, sources, model behaviour and downstream actions can go wrong, and build cases for each one.
Input failures include missing identifiers, contradictory instructions, unsupported file types, very long material and languages the feature wasn't built for. Source failures include an unavailable index, stale content, duplicate versions, a document outside the user's permissions, and a source that simply doesn't contain the answer.
The model might ignore a constraint, follow an instruction buried in retrieved content, invent a source, produce invalid structure, repeat sensitive input back, or pick the wrong tool. Tool calls bring their own problems: timeouts, partial responses, duplicate submissions, and success messages that turn up after the application has already marked the action as failed.
For each failure mode, decide what containment you want. The workflow might stop before taking an action, show a clear error, ask for approval, keep a draft, or put the case into a review queue. The expected behaviour should cover state as well. A friendly error message doesn't help much if the system has already changed a record and can't tell you whether that change went through.
Combine failures too. A missing source plus a leading user request can produce a confident unsupported answer. A tool timeout followed by an automatic retry can create a duplicate action. A long input can push an important instruction out of the model's working context. Checks on individual components tend to miss these paths.
If the feature can read private material or take actions, security testing belongs in this set. Try asking for information outside the test identity's access. Put hostile instructions inside a document the user is allowed to see. Try to get the workflow to reveal its hidden configuration or use a tool outside the task boundary. Then check what actually happened in the target system or the permission log, instead of taking the assistant's word that it refused.
Keep a failure register next to the evaluation results, with the condition, what you observed, severity, how it was detected, the current control and the retest case. You'll probably accept some failures for a controlled release, and each of those needs an owner and a visible plan for what happens when it occurs.