20 March 2025 / AI agents

Scoping tool access for an AI agent

OpenAI's Responses API combines models with search, files and computer use. Each tool gives the agent a different kind of authority, so I define its identity and approval rules before connecting it.

Circular openings and railings in a pale concrete building.
Photo: Joel Filipe (opens in a new tab)

OpenAI's Responses API combines models with search, files and computer use. Each tool gives the agent a different kind of authority, so I define its identity and approval rules before connecting it. OpenAI released the Responses API on 11 March with built-in web search, file search and computer use. The announcement also introduced an Agents SDK with guardrails and tracing. Applications can combine those tools in one response workflow. The application team still has to define permissions. "The agent can use search" says little about what it can see, change or expose. I do not want an agent borrowing an administrator's session or a developer's personal token. It needs a named service identity with permissions tied to a particular job. That identity should answer practical questions. Which environment can it reach? Which file collections can it search? Can it read records for every customer or only the current account? Can it create a draft, or can it publish? If the credential appears in a log, how quickly can it be revoked? Use the narrowest credential the tool supports. Prefer short-lived tokens over permanent keys. Keep production and test identities separate, and make the production identity visibly different so a local experiment cannot drift into live data. If a connector cannot express the required scope, that is a property of the connector, not a reason to grant broad access and hope the prompt behaves.

The model should never receive secrets it does not need. Tool execution can happen behind a trusted boundary where the application validates arguments, adds the credential and filters the response before returning it to the model. Search gives an agent fresh information, but the returned pages were written by somebody else. Their contents can be wrong, malicious or irrelevant. A page can also contain text that looks like an instruction to the model. Preserve source URLs and retrieval times with search results. Keep page content clearly separated from system and developer instructions. For a material claim, require the final response to cite the source actually used so a reviewer can open it. I would also restrict where search results can flow. Public web content should not be able to trigger an internal action without another validation step. If the agent finds a bank account number, package command or configuration value online, it can report the finding. It should not execute or store it merely because it appeared in a result. Search queries themselves may disclose context. Strip private names, identifiers and document fragments before sending a query to an external search service. In some workflows the safe setting is no web search at all.

File search sounds read-only, but reading can still cross a confidentiality boundary. A shared vector store containing documents from several accounts may retrieve text from the wrong one even when the query never names it. Partition collections according to the product's access model. Attach tenant or project identifiers as enforced filters, then test that a query from one scope cannot retrieve another scope's chunks. Do not rely on the prompt to say "only use files for this customer". The retrieval layer needs to enforce it. Ingestion also needs controls. Record the source document, version, access classification and ingestion time for every chunk. When a document is withdrawn or a person's access changes, there must be a deletion path that reaches the search index, not only the original file store. Return the document and passage reference with retrieved text so the agent can point back to its source. If several versions conflict, surface the conflict rather than blending them into one confident answer. The computer-use tool translates model output into mouse and keyboard actions. OpenAI describes it as a research preview and recommends isolation plus user confirmation for consequential actions. That is a sensible starting position. Run browser or desktop automation in a sandbox with an explicit allowlist of destinations. Start with a clean profile rather than a person's normal browser, which may contain active sessions, saved passwords and unrelated tabs. Disable downloads or uploads unless the job needs them. Put file-system and network restrictions around the sandbox as well.

The action policy should distinguish low-consequence navigation from commitments. Opening a page and reading a field may run automatically. Sending a message, submitting a form, purchasing something, changing permissions or deleting a record should pause with a plain preview of the proposed action. Confirmation has to be specific. "Continue?" is weak. Show the destination, affected account, values being submitted and any irreversible consequence. After approval, bind it to that exact action. If the page changes or the values differ, ask again. Schema-valid JSON does not prove that a tool call is authorised. The executor should validate every argument against the current identity and scope. For a write operation, check the authenticated identity, resource scope, permitted fields and current state. Apply size and rate limits. Reject paths, URLs or identifiers outside the allowlist. Where a tool supports both reads and writes, expose them as separate functions so the agent cannot reach a write through an ambiguous parameter. I also want a dry-run mode for actions that can be represented before execution. The tool can return the proposed database update, email or command without applying it. The agent and reviewer then work from the same concrete payload.

Tool errors must be safe to show to the model. Raw exceptions can include credentials, internal paths or data from another scope. Return a scrubbed error code and enough detail to decide whether retrying makes sense. A useful trace records the boundaries that affected each action. For each tool call, retain the agent identity, tool version, validated arguments, target, approval event, result status and source references. Sensitive payloads may need redaction or controlled storage, but the record should still explain the action. Before release, test the denied paths: a search containing private data, a cross-project file query, a write outside the allowed resource and a changed computer-use action after approval. Confirm that the system blocks each one and that the trace says why. Start the first production workflow with one identity, a small tool allowlist and explicit review on every write. Expand a permission only when a real task needs it and the logs show the current boundary working.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs