25 September 2025 / Applied AI / 9 chapters

Which agent decisions are worth capturing

From Designing an audit trail for tool-using agents

Token-by-token logs cost a lot to keep, and they still won't tell you which source, policy result or tool event authorised an action. I'd capture explicit decision records at the points where the workflow branches or where authority changes hands.

Useful events here are a plan accepted for execution, a tool selected, a policy check, a proposal created, an approval requested, an action attempted and the task outcome. Each one should carry the relevant structured payload, or a reference to it. You don't want a model-written story about what happened.

For a tool selection, record the tool name and version, the validated arguments with secrets removed, the requesting run, the policy result and a reason code. If the gateway changes or fills in an argument, keep both what was requested and what actually ran. The model might ask for one date range and policy might narrow it to another.

For model calls, record the provider and model identifier, request time, response classification, token or cost data if you have it, and links to the context artefacts. You don't need full chain-of-thought for an operational audit trail. It might not be available anyway, and generated reasoning text doesn't replace the actual source, policy and tool events.

If the workflow produces a rationale for a proposal, store it as model output and label it that way. Your evidence is still the cited records and the deterministic checks that passed. A plausible rationale can help a reviewer and still be wrong.

Denied actions go in the trail too. They show enforcement worked, and they can show the agent repeatedly trying to reach the wrong tool, object or scope. I'd use stable denial codes like scope_mismatch, restricted_field, approval_missing, stale_version and limit_reached so you can query for patterns.

Retries should link back to the original call and say why another attempt was allowed. Without that, a few near-identical action events can look like duplicate writes, or hide the fact that the first outcome was uncertain.

All articles