Blog
What computer use exposes in business software
Anthropic's computer-use beta can click, type and navigate software without a purpose-built API. That makes every ambiguous label, hidden state and irreversible button part of the agent's operating environment.

Anthropic released its computer-use capability in public beta this month. The company describes it as experimental and at times cumbersome and error-prone. That warning matters. A model operating a screen inherits all the uncertainty people work around, while lacking much of the context a regular operator has built up. With computer use, interface labels, visible state and controls affect what the model does. Inspect those screens and restrict the account before giving an agent a login. People learn that "Process" on one screen means preview and on another means commit. They remember which icon opens a menu and which grey control is disabled. A computer-using model has to infer those meanings from the current screen and its instructions. Ambiguous labels increase the chance of selecting the wrong control. Use verbs that name the effect: "Create draft invoice", "Send message" or "Delete import". If an action has scope, show it in the control or nearby text. "Archive 126 records" is more informative than "Continue". Controls also need stable accessible names and clear visual states. An icon-only button, a label that changes position or two adjacent buttons with similar appearance can make operation brittle. This is an accessibility problem for people too, which is a useful reason to improve it regardless of agent plans.
Test at realistic screen sizes and zoom levels. A control outside the viewport or hidden behind a sticky panel may not exist from the agent's current view. Business software carries state in selected filters, active accounts, unsaved edits, background jobs and permissions. A person may notice the customer name in a header or remember that they opened a second tab. An agent can miss that context and perform a valid action in the wrong place. Put consequential state close to the action. Before a bulk update, show the account, filter, number of affected records and proposed change together. Avoid relying on colour alone. If a previous step changed the working context, repeat that context on the confirmation screen. Loading and saving states need explicit treatment. Disable a control while a request is still running, show whether a save completed and prevent duplicate submission. A spinner that disappears without a result leaves both a person and an agent guessing whether to retry. Sessions can expire during a long task. Redirecting silently to a login page may cause the agent to type task data into an email field. Detect expiry, stop the workflow and require re-authentication through a controlled path. A computer-use agent can reach any control its account and interface expose. Restrict the account before trying to solve everything with prompting. Give it only the applications, records and actions needed for the named task.
Separate preparation from execution. The agent can fill a draft, assemble a batch or navigate to the final step, while a person approves sending, deleting, paying or publishing. The approval should display the proposed effect, not a generic question asking whether to proceed. Some actions should remain unavailable to this route. Credential changes, permission grants and exports of sensitive data may require a different identity or a direct administrative process. If a task cannot be completed with the restricted account, that is useful information about the proposed automation. For each allowed action, record whether it can be reversed, how long reversal remains possible and what support will need to trace it. When an agent reads web pages, messages or documents and can also take actions, untrusted content may try to redirect its behaviour. Text that looks like an instruction to the model can appear inside the very material it has been asked to process. Treat screen content as data unless the workflow explicitly identifies it as an authorised instruction source. Keep the task instruction separate from content being reviewed. Restrict navigation to approved domains and block downloads, credential entry or external communication that the task does not require.
A person should approve changes that cross a trust boundary. That includes sending information to a new destination, opening an unexpected login page or following a link outside the application. The agent should stop when the interface diverges from the expected route rather than improvise through it. Use synthetic data and isolated accounts during testing. Do not begin by turning an experimental operator loose inside a live administrative session. A chat transcript does not fully describe what happened on screen. An operational record should connect the task to navigation steps, observations, entered values, clicks, resulting state and approvals. Capture screenshots or structured observations at important boundaries, while applying the same privacy controls used for other sensitive logs. Record the application, account identity, time, action and result. If the agent retries, retain that sequence rather than replacing it with the successful final attempt. The log should answer practical questions: Which record was open? What did the agent believe it was doing? Which control did it select? What changed? Who approved the final action? Could the same step be run again safely? Do not record passwords, session tokens or unnecessary personal information. Observability that creates a new credential store is a poor trade.
A useful computer-use test is more than a video of one successful path. Write the expected states and allowed transitions, then interrupt the path. Test a slow page, validation error, expired session, changed layout, unexpected modal, empty result and duplicate click. Change a filter halfway through. Remove a permission. Put an untrusted instruction into displayed content. At every interruption, check whether the agent stops, recovers safely or makes the situation worse. Start in read-only work where the output can be compared with the screen. Then allow draft creation. Add an approval boundary before any external or destructive action. Keep a manual stop control outside the interface being operated, so it remains available if navigation goes wrong. Choose one proposed workflow and mark every screen state the agent must understand. At each step, check that the screen identifies the current object, pending work, action scope and recovery route. Missing information should be fixed in the interface before the trial can take actions in that step.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.