25 September 2025 / Applied AI / 9 chapters

Recording inputs without copying everything

From Designing an audit trail for tool-using agents

You need the audit trail to show what context the agent used. If you copy every prompt, document and private record into a general logging platform to do that, though, you've just built a second content store that nobody's controlling.

For the task request, I'd store a redacted or access-controlled version of the user's instruction, plus a hash of the exact original if you need integrity checks. Attachments get recorded by object ID, version, checksum, classification and the bounded portion you extracted. The original stays where it already lives, under the controls it already has.

System instructions and tool definitions should be versioned artefacts. An event can then carry instruction_version, tool_registry_version, the model identifier and whichever runtime settings matter. If instructions are assembled on the fly, keep the component versions and a hash of the rendered input. Only store the full rendered prompt where policy allows it and you've got a genuine review case that needs it.

Retrieval events should say which source objects came back, their versions, the access decision and whether results were truncated. If a document only contributed two passages, record where those passages sit instead of copying the whole file.

Make it obvious where a source came from. A current system record, an uploaded document and an unverified web page shouldn't look the same in the log. The event can carry source type, owner, effective date and retrieval method, and you don't have to ask the model to describe any of that later.

Mark untrusted content on input events as well. Tool outputs, files and web pages can contain text that looks like an instruction, and having that classification recorded lets a reviewer tell the difference between instructions someone authorised and material the agent just happened to read.

When an input couldn't be loaded (permissions, a timeout, a parse failure), log that. A final answer can look complete even though a source named in the task never made it into the context.

All articles