Article chapter 06 of 08
Keep the records support will need
Logs need to answer a support or review question without dumping every private input into one place. Design the event record around the workflow states and decisions. Each event should have a run identifier, event type, timestamp, actor or service identity, workflow version, target reference, result and correlation identifier for any external request.
Keep raw model input and output only where the purpose, access and retention period are approved. In many workflows, structured references and a redacted output provide enough routine visibility, with authorised access to fuller evidence for investigation. Logging everything by default can create a second, poorly governed store of sensitive material.
Operational monitoring should include runs waiting for review, unknown external actions, source ingestion failures, permission denials, validation failures, tool timeouts and the age of the oldest unresolved item. Model latency by itself will not show that a queue has stopped moving.
Alerts need an owner and a response. Write down which condition pages somebody, which creates a work item and which is reviewed on a schedule. Include the link or query that finds affected runs. An alert saying "AI error rate increased" is hard to act on if it does not identify the task family, configuration or failed stage.
Build a support view before the first broad release. It should find a run by customer-safe reference, show its current state, list approved actions and display source and tool evidence according to the operator's access. The view should avoid exposing hidden prompts or unrelated private data simply because support needs to correct one task.
Exercise the evidence trail with a prepared failure. Ask somebody who did not build the feature to determine what happened and propose the safe next action. Missing timestamps, ambiguous state names and inaccessible source links become apparent quickly. Update the runbook and interface from that review.