Blog
Where agent supervision has to happen
Anthropic's recent analysis suggests experienced users approve more automation while also interrupting agents more often. That points to supervision built around visible plans and consequential boundaries instead of a confirmation attached to every command.

Anthropic's recent analysis suggests experienced users approve more automation while also interrupting agents more often. That points to supervision built around visible plans and consequential boundaries instead of a confirmation attached to every command. In Measuring AI agent autonomy in practice, published on 18 February, Anthropic reports that newer Claude Code users employ full auto-approve in roughly 20 per cent of sessions, rising to more than 40 per cent among experienced users. Interruptions also rise, from about 5 per cent of turns for newer users to around 9 per cent for more experienced users. Those figures come from one coding product, and Anthropic notes several limitations. They are still useful for thinking about interaction design. Approval and supervision are not the same event. Per-command approval is understandable when an agent first gains access to a shell, file system or external service. The user sees each proposed action and decides whether it may proceed. That works for short tasks with a few meaningful steps. It degrades on work that needs dozens of reads, edits and test runs. The user starts approving familiar commands based on shape rather than purpose. A command may be safe in one directory and damaging in another. A package installation may be expected during setup and suspicious halfway through a documentation task.
The number of prompts also says little about whether the user understood the plan. Someone can approve twenty commands without noticing that the agent is solving the wrong problem. Keep individual approval for actions where the specific parameters carry the risk. Sending a message, changing production data, publishing a release, rotating a credential or running a destructive migration deserves a final check tied to the actual target. Routine reads inside an approved workspace can usually be governed at a broader level. A visible plan gives the user a chance to correct intent while the cost is low. It should name the objective, proposed stages, expected files or systems, verification and any action that will require another decision. The plan does not need to predict every command. Agents encounter missing dependencies and failing tests. It should provide enough structure to tell whether a new action belongs to the agreed job. Approval of the plan should create a bounded execution envelope. For example, the agent may read the repository, edit one application package and run local tests. Network access, changes to deployment configuration or writes to an external service remain outside that envelope. When the plan changes materially, show the difference. "One test failed" is ordinary execution feedback. "The fix now requires changing the shared authentication library" is a new scope decision. The interface should pause there, explain why and let the user accept, redirect or stop.
Interruption only works when the user can see enough to judge the run. A spinner and elapsed time are poor supervision tools. Show the current stage, recent meaningful action, files or resources changed and verification status. Raise an alert when the agent moves outside expected scope, repeats the same failed action, requests broader permissions, encounters ambiguous data or prepares an irreversible operation. Do not stream every token into the main supervision view. Detailed logs should remain available, but the active view needs a small set of stable signals. Otherwise the user is technically informed and practically unable to follow what matters. The stop control must be immediate and honest. It should say whether it cancels future tool calls, terminates a local process, rolls back uncommitted work or merely asks the model to stop after its current action. Those are different guarantees. After an interruption, preserve the partial state. The user should be able to inspect what changed, add direction and resume, or discard the run without guessing which commands already completed. Supervision should become stricter as consequences increase. A useful tool definition includes the identity the agent uses, resources it can reach, allowed operations, validation rules and whether the action can be reversed.
Separate read and write tools where possible. Reading a deployment status and starting a deployment should not share one broad permission. Use short-lived credentials scoped to the task rather than a developer's full session. Require stable identifiers in write calls so the model does not choose a target from a display name. For important writes, present the proposed action in business terms before execution: which record, environment or recipient; what will change; and whether a recovery path exists. A raw command may still be available for technical review, but it should not be the only explanation. The agent also needs a clear way to decline action when required information is missing. Guessing an account, environment or date range to avoid another prompt is poor autonomy. A successful tool response proves that the tool accepted a request. It does not prove that the intended result exists. After a state-changing action, read the target back through an independent route where practical. After code changes, run the relevant checks and inspect the diff. After an integration write, confirm the destination record and correlation identifier. For a deployment, check the released revision and health signal rather than relying on the command exit code.
The supervision interface should distinguish attempted, accepted and verified. That distinction helps the user decide whether to interrupt a run that appears busy but is no longer making progress. Verification also needs a failure route. If the read-back disagrees, the agent should stop further dependent actions and show the mismatch. Continuing from an unverified state can turn one small error into a chain of plausible but invalid work. A complete command transcript is useful during diagnosis, but it is difficult to review as a work record. Keep the plan that was approved, later scope changes, permissions granted, consequential actions, interruptions, verification results and final disposition. Record who approved each boundary and which identity executed it. Link summaries to the underlying tool events so a reviewer can inspect detail without treating generated prose as evidence. Test the design with one long but low-risk task and one short consequential action. The first should run without constant prompts while remaining easy to redirect. The second should stop at the exact point where target and effect are known, then prove what happened afterwards. If both tasks produce the same approval pattern, the supervision model is probably attached to commands rather than consequences.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.