Article chapter 09 of 09
Test it with one clean run and one broken one
Before release, set up a production-like task that reads controlled private records, produces a proposal, gets approved and makes a reversible external change. Follow it through every event and compare your audit view with the source system's logs.
Check that the trail has the task and run IDs, the requesting actor, the workload credential, instruction and tool versions, source object versions, redacted tool arguments, policy results, the proposal hash, the approval, the executed payload and the destination receipt.
Then break a second run on purpose: interrupt it after the action request has gone out but before the local acknowledgement is stored. The task should stay unresolved. Reconcile it through the destination, record the outcome, and check that the recovery view only clears once that evidence is in.
Run the edge cases as well: a denied cross-scope read, a truncated search, a stale target version, an expired approval, an edited proposal, a replayed action ID, a secret inside a tool error and deletion of retained prompt content. Look at what the system records, and also at what it refuses to record.
Last, hand it to an operator who didn't build the workflow and ask them to answer the questions from the first section using only the authorised audit views and the linked source records. If any answer needs the developer to remember something, you're missing an event, a status is unclear, or one of the views needs more work.
When a run goes wrong and the final chat message has squashed it into a few lines, there are questions the transcript can't answer. The audit trail I'd design is what lets you answer those questions without leaning on the transcript: one task ID across retries, inputs recorded by reference, approvals tied to an exact payload hash, and outcomes that stay uncertain until the destination confirms them. If you're adding tools to an agent now, I'd start with the task and run IDs and the action event that gets written before each external call, then run the interrupted-action test above before anyone relies on it.