17 October 2024 / Applied AI / 8 chapters

Exercise failure modes deliberately

From Evaluating an AI workflow before customers depend on it

Waiting for an evaluation set to encounter failures by chance leaves predictable gaps. Use the workflow map to identify where inputs, sources, model behaviour and downstream actions can fail, then construct cases for each path.

Input failures include missing identifiers, contradictory instructions, unsupported file types, very long material and language the feature was not designed to handle. Source failures include an unavailable index, stale content, duplicate versions, a document outside the user's permission and a source that does not contain the requested answer.

The model may ignore a constraint, follow an instruction embedded in retrieved content, invent a source, produce invalid structure, repeat sensitive input, or select the wrong tool. Tool calls add their own failures: timeouts, partial responses, duplicate submissions and success messages that arrive after the application has already marked the action as failed.

For each failure mode, define the desired containment. The workflow might stop before an action, show a clear error, request approval, preserve a draft, or place the case into a review queue. The expected behaviour should include state. A friendly error message is insufficient if the system has already changed a record and cannot say whether that change completed.

Combine failure conditions as well. A missing source and a suggestive user request may produce a confident unsupported answer. A tool timeout followed by an automatic retry may create a duplicate action. A long input may push an important instruction out of the model's working context. These cases exercise paths that individual component checks miss.

Security testing belongs in this set where the feature can read private material or take action. Try requests for information outside the test identity's access. Put hostile instructions inside a permitted document. Attempt to make the workflow reveal hidden configuration or use a tool beyond the task boundary. Verify the result through the target system or permission log, rather than accepting the assistant's statement that it refused.

Keep a failure register beside the evaluation results. Record the condition, observed behaviour, severity, detection method, current control and retest case. Some failures will be accepted for a controlled release, but they need an owner and a visible operating response.

All articles