Article chapter 08 of 08
Test the boundary before you release it
I'd start with a permission matrix. Put the task steps down one side and the source or action scopes across the other, and for every allowed cell write down why it's needed. If you can't write a reason, that permission is a candidate for removal.
Then test with the real workload identity. Cover the expected work, plus some deliberate attempts to push past the boundary:
- request a record from another tenant
- ask for a restricted field on an allowed record
- include a conflicting tenant ID in the tool arguments
- put tool-like instructions inside an attachment
- request more results than the configured limit
- try a tool name that isn't registered
- change a proposal after it's been approved
- replay an approval that's already been used
- change the target record before execution
- interrupt an external action after it's sent but before it's acknowledged
For each one, check that the system denies the call or routes it to the review path you intended. Look at what the user is told and at the event record. A denial that's technically correct but still reveals that the restricted record exists, or leaks what's in it, needs fixing.
Run the approved path too. Check that the agent pulls only the fields it needs, cites the source objects it used, builds a complete proposal, waits for the right reviewer, executes once and returns evidence from the destination. If a permission didn't get used anywhere in that run, take it out.
The release record should name the task version, identity, source scopes, tool set, policy version, approval rules, operating limits, the failures you tested and the owner. When the next task comes along, start from that record and add only the permissions the new task needs.
That's the trouble with a request like "help with customer accounts". Deciding what a tool-using agent can read and change comes down to tying every permission to one written task, so each read, proposal and change can be explained by a step in that task and enforced outside the model. If you've been handed a request like that, I'd write the first task down as something you could watch happen, fill in the permission matrix for it and take out any access you can't give a reason for before the agent runs under its own identity.