19 December 2024 / Applied AI / 8 chapters

Make the runtime predictable enough to operate

From Production readiness for applied AI

The generative step will vary, but the runtime around it can still have explicit limits and state. Set request timeouts, tool timeouts, retry rules, maximum attempts, input limits and output limits, and decide what the user sees when each one is hit.

A failed model request can usually be tried again if nothing happened downstream. A tool action is different: it needs an idempotency key or a reconciliation check before you retry it. Keep retries bounded and record every attempt under the original run. Endless retry loops waste capacity and can leave work stuck somewhere nobody's looking.

Validate model output before it reaches business logic. If the workflow expects structured data, enforce the schema and field types. Check identifiers against allowed values, and recalculate deterministic things like amounts or dates in application code. If the output includes actions nobody asked for, reject them, even if the explanation that comes with them reads well.

Keep proposal and execution separate. Store the model's proposed action, run it through deterministic validation and policy checks, then get approval where it's required. The execution component should only ever receive a fixed, validated operation. That way you can test the model, the policy and the connector paths on their own.

Where two runs can touch the same record, control concurrency. A version check, a lock or a compare-and-set rule stops an old proposal from overwriting a newer change. Show the conflict to the reviewer along with the current state. If you quietly rerun the model against the changed data, the original approval doesn't mean anything any more.

For capacity planning, start from what you actually saw in testing. Record input size, output size, tool calls, latency, how often things time out and how fast the reviewer queue grows during controlled use. Set service limits that protect other workloads, and when a limit is reached, reject or queue the work visibly instead of accepting requests that might never finish.

Test degraded operation too. Turn off retrieval, return a malformed tool response, slow down an external system and run through a configured quota. Does the user get an accurate status? Can the run be resumed or safely abandoned? That tells you whether operational state exists outside the happy path.

All articles