10 September 2026 / Agent operations

What stays yours when the agent harness is managed

OpenAI's Agents API now provides hosted sessions, environments, context management, tool loading and subagents. Product teams still have to define the work, authority, evidence and recovery path around that managed harness.

An empty industrial test bay with a steel workbench beneath a red overhead gantry.
Generated image: Villar David editorial

OpenAI's Agents API now provides hosted sessions, environments, context management, tool loading and subagents. Product teams still have to define the work, authority, evidence and recovery path around that managed harness. The Agents API announcement describes a public beta for agents that can keep working across long sessions, run in managed or partner environments, find tools as needed and delegate work to subagents. This removes a substantial amount of runtime engineering for a team that would otherwise build the harness itself. It also makes the boundary between platform responsibility and product responsibility worth writing down. The first thing that stays with the product is the definition of the job. "Investigate the issue" may be enough for a person who already knows the system and the expected report. An agent needs the target, relevant time window, permitted sources, output location and a stopping condition. If it is allowed to change something, it also needs the exact change boundary and the evidence required afterwards.

The environment is another product decision even when someone else hosts it. Decide which files enter the session, what can leave it and what survives after the run. A clean sandbox can isolate execution, but it does not decide whether a customer export belonged there. The application needs to make that decision before the task starts and retain enough of the input record to explain it later. Tool access should be assembled for the task rather than inherited from the developer who created the integration. A long-running agent that can search the web, read a repository and query observability data has three different trust boundaries. If the work also needs a deployment tool, I would add that only to the stage that has passed review. Keeping every possible tool available for the whole session makes the task easier to describe and harder to govern.

Managed context does not remove the need to choose what the agent should remember. OpenAI says the API can compact earlier context as a session approaches its limit and load tool definitions on demand. Those mechanisms help the harness keep working. The product still needs durable state for decisions that must not depend on a summary: the approved input version, operation identifiers, external results, review status and unresolved questions. Subagents introduce the same issue at another level. Parallel work is useful when the pieces are independent. The main agent needs to know what each subagent was asked to do, which tools it received and what evidence came back. A consolidated answer without those links is difficult to review because a confident conclusion can hide one failed or incomplete branch.

I would keep the activity record outside the agent's own final response. Record session and task identifiers, model and harness version, tools granted, important external actions, artefacts produced and the checks run before completion. The agent can write a useful summary, but the application should collect the facts that establish what actually happened. Recovery is part of this contract. A session can stop after an external action succeeds and before its result is stored locally. A subagent can finish while the parent fails to collect its output. A tool can time out after accepting a request. The workflow needs states for those uncertain outcomes and a way to reconcile them before another run repeats the action. Cost controls also belong near the work definition. A managed service can expose model, tool and environment usage, but a product owner still has to decide what a completed task is worth. Set a time or spend boundary, retain partial artefacts when that boundary is reached and distinguish a stopped task from a failed one. Otherwise the team will learn about an open-ended job from the bill or an operator waiting for a result.

I would start with one workflow whose result a reviewer can check independently. Run it through the managed harness, then inspect the record without relying on the final answer. Can the reviewer see what the agent read, which actions reached external systems, what remains unresolved and where to resume? If those questions are hard to answer, the missing work sits in the product around the harness.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs