Blog
Govern the whole agent system
I review an agent by tracing what it can read, what it can change and how a person can stop or reverse it. That includes the harness and credentials as well as the model.

I review an agent by tracing what it can read, what it can change and how a person can stop or reverse it. That includes the harness and credentials as well as the model. A model might generate the reasoning, but it does not decide on its own which database is connected, which tools are exposed or whether a command runs immediately. Those decisions sit around the model. They are spread across prompts, tool definitions, service accounts, job runners, approval screens and ordinary application code. Reviewing only the model leaves most of the operating system out of view. A useful review starts with a named job, described in the same terms a person doing the work would understand. "Help with operations" is too broad. "Read a failed import report, identify the affected batch and prepare a retry for approval" gives the review something concrete to follow. For that job, I want to know where the request enters the system, what information is loaded, which tools become available and what the agent can return. This exposes decisions that are otherwise hidden inside configuration. An agent may have read-only access in one tool and write access through another. A harmless-looking file tool may reach a directory containing credentials or customer exports. A retry action may also trigger notifications or downstream processing.
The scope should be written down before testing starts. That makes later changes visible. If a new tool is added, or a read action becomes a write action, the job definition and review need to change with it. An agent acts through identities created by the surrounding software. Sometimes it uses one service account for every user. Sometimes it receives a short-lived token derived from the person who started the task. Those arrangements produce different permissions, logs and recovery problems. For each connected system, I record:
- which identity the tool uses;
- how the credential is issued and stored;
- which records or actions it permits;
- whether access differs by user, role or environment;
- where an operator can revoke it.
This is ordinary access-control work, but agent interfaces make it easier to miss. A chat window can look personal while every request behind it runs as the same powerful account. The user may assume the agent can only see what they can see. The integration may say otherwise. Credentials also outlive prompts. Changing the system instruction does not remove a token from a worker, close a network path or invalidate a queued job. Permission changes have to happen in the systems that enforce them. A confirmation prompt on every tool call creates noise and teaches people to approve without reading. I place checks where the consequence changes: sending a message, modifying a record, moving money, publishing content, deleting data or starting work that is expensive to unwind. The approval should show the proposed action in terms the reviewer can recognise. A raw function name and a JSON payload may be useful in the activity log, but they are a poor approval screen. The reviewer needs the target, the intended change, relevant side effects and whether the action can be reversed. Some actions can run automatically within a narrow boundary. Reading a permitted document set or drafting a change in an isolated workspace may be acceptable without repeated interruption. The boundary still needs enforcement outside the prompt. Tool code can restrict paths, validate identifiers, cap batch sizes and reject unsupported operations before execution.
The harness chooses what context reaches the model and how its output becomes a tool call. It may summarise earlier messages, inject stored memory, retry failed requests or continue a task after the user has left. Each behaviour can alter what the agent knows and what it does. The job runner adds another layer. A person may stop the visible conversation while a background process keeps working. Queued tasks may start with older permissions or an outdated version of the instructions. Retries can repeat an action that succeeded but failed to report its result. I check how the system records a job identifier, instruction version, tool calls, approvals and final state. I also check what happens after a timeout or partial failure. If the system cannot tell whether an external write completed, it should stop for reconciliation rather than issue the same write again. A stop button needs a defined reach. Does it cancel only the next model request, or does it also terminate running commands, revoke queued work and prevent a retry worker from starting again? The answer should be tested against the actual components.
Recovery needs the same treatment. For reversible changes, the interface should retain the information needed to undo them. For irreversible actions, the system should require stronger review before execution and leave a clear record afterwards. Restoring a database backup is not a practical undo mechanism for one badly scoped agent action. I would test at least a cancelled job, a tool timeout, an expired credential, a duplicated request and a result that reaches one system but not another. These tests reveal whether the operator can see the incomplete state and choose a safe next step. An agent review should run again when its tools, permissions, memory sources, approval rules or execution environment change. A model update may justify new evaluation, but a small integration change can expand authority just as easily. The practical artefact is a short trace for each approved job: entry point, context sources, identities, tools, consequential actions, stop mechanism and recovery path. Keep it close to the implementation and test the paths it describes. When one of those entries changes, the review has a precise place to start.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.