19 December 2024 / Applied AI / 8 chapters

Make runtime behaviour predictable enough to operate

From Production readiness for applied AI

Give the runtime explicit limits and state even though the generative step can vary. Set request timeouts, tool timeouts, retry rules, maximum attempts, input limits and output limits. Decide what the user sees when each limit is reached.

Retries require care. A failed model request can usually be attempted again if no downstream action occurred. A tool action needs an idempotency key or a reconciliation check before retry. Keep retries bounded and record each attempt under the original run. Endless retry loops waste capacity and can hold work in a state nobody sees.

Validate model output before it reaches business logic. If the workflow expects structured data, enforce the schema and field types. Check identifiers against allowed values and recalculate deterministic amounts or dates in application code. Reject extra actions that were never requested, even when the accompanying explanation reads well.

Separate proposal from execution. Store the model's proposed action, pass it through deterministic validation and policy checks, then seek approval where required. The execution component should receive a fixed, validated operation. The model, policy and connector paths can then be tested independently.

Control concurrency where two runs can affect the same record. Use a version check, lock or compare-and-set rule so an old proposal cannot overwrite a newer change. Display the conflict to the reviewer with the current state. Quietly rerunning the model against changed data makes the original approval meaningless.

Begin capacity planning with observed test behaviour. Record input size, output size, tool calls, latency, timeout frequency and reviewer queue growth during controlled use. Set service limits that protect other workloads. When a limit is reached, reject or queue work visibly instead of accepting requests that may never finish.

Test degraded operation. Disable retrieval, return a malformed tool response, slow an external system and exhaust a configured quota. Check whether the user receives an accurate status and whether the run can be resumed or safely abandoned. These tests reveal whether operational state exists outside the happy path.

All articles