Article chapter 01 of 09
What someone will ask you after a run goes wrong
A tool-using agent gets a task, pulls together context, picks tools and sometimes asks a person to approve an action. Then it writes a final chat message that squashes all of that into a few lines. That message can leave out denied calls, retries, the sources it read along the way, proposals that got edited, and an external action that finished after the model had already stopped responding.
So when a result looks wrong, the transcript usually can't answer the questions people actually ask. Things like:
- Who started the task, and what identity did it run under?
- What was the objective, and what limits applied?
- Which private objects did the agent read?
- Which model, instructions and tool definitions were in play?
- What did the agent propose, and what exactly did someone approve?
- Which calls reached an external system?
- What did each destination confirm?
- Is anything still uncertain or half done?
- Can you work out the sequence without opening sensitive content you don't need to see?
I'd answer those from structured events. The transcript is still worth keeping, but as one source sitting next to tool events, policy decisions, approvals and destination receipts.
It's also worth thinking about who'll be looking. An operator recovering a failed task wants a short sequence and the current status. A security reviewer might want denied access attempts and identity details. A user might just want a plain account of which records were viewed and what was changed. They can all read from the same event stream, through different views with different access rules.