Article chapter 06 of 08
Sort actions by what happens if they go wrong
The HTTP method doesn't tell you much about the permission an action needs. Sending an email, publishing a file, creating a user, moving money, changing access and kicking off a long-running job all have different consequences and different ways of recovering, so I'd classify actions on those properties.
The questions I'd ask for each action are pretty concrete. Who can be affected? Can it be undone? Does undoing it create another external event? Is there a deadline? Can the destination deduplicate a repeated request? Will another system act on the result automatically? Does it expose private information to someone new?
Policy rules can then require stronger review for particular values or destinations. Say the agent drafts a message to an address that's already attached to the case. That might go through the normal approval path, while a brand new external recipient needs another check. Keep those rules in the tool gateway or a policy service. Wording in the prompt is useful context for the model, but it can't enforce anything.
Execution has to be idempotent. Give each approved proposal a stable action ID and pass it as an idempotency key wherever the destination supports one. Record the destination's response and its object identifier. If a timeout leaves you unsure whether the action went through, query by that key or send the task off for reconciliation before you retry.
After execution, return evidence: the fields the destination accepted, its version or reference, the completion time and any downstream work that started. A tool call returning without an exception doesn't mean the action succeeded, and a queued action can still fail later.
Reversal should be a named operation with its own permission. I'd avoid a generic undo button that has to guess how to compensate for several different kinds of side effect. If an action can't be reversed, show that before approval and keep the permission for it as narrow as you can.