Blog
Production agents need institutional context
OpenAI's internal data-agent write-up shows how much context sits around the model: schemas, code, permissions and previous corrections. The useful design question is who maintains that context and how a user can correct it.

OpenAI's internal data-agent write-up shows how much context sits around the model: schemas, code, permissions and previous corrections. The useful design question is who maintains that context and how a user can correct it. OpenAI published Inside OpenAI's in-house data agent on 29 January. The interesting part is the surrounding system. The agent draws on metadata, table lineage, historical queries, human annotations, code, internal documents, remembered corrections and live inspection. Access follows the user's existing permissions. That is a lot of maintained context for a product that appears to accept a natural-language question. Most organisations considering an internal agent have some of these ingredients, but they are split across systems and owned by different people. A model can infer the shape of a table from its columns. It cannot reliably know that a field excludes a class of transactions, that a metric changed definition last quarter, or that two similar datasets are intended for different teams. Those details live in code, operating notes and people's corrections. Putting all available text into retrieval will not resolve the problem. The agent needs context selected for the task, with enough metadata to judge authority and freshness. A useful context record should say what it describes, where it came from, who owns it, when it was last checked and which users may retrieve it. Ownership is the uncomfortable field. A data team may own schemas while a finance team owns the meaning of a metric. Engineering may own the code that derives it. If nobody is accountable for the combined description, the agent can retrieve a technically current definition that is operationally wrong.
Treat the context catalogue like another production dependency. Give it health checks, review dates and a visible path for reporting a bad record. Schemas describe available fields and types. They rarely explain all the transformations that produced a dataset. A column called active_customer may depend on exclusions, time windows and upstream events that are only visible in pipeline code. OpenAI describes using code-level definitions to enrich its agent's understanding of tables. That suggests a practical pattern beyond analytics. An internal agent answering questions about software should be able to connect a user-facing concept to the code, configuration and jobs that implement it. The connection should be traceable. If the agent states that a figure refreshes daily, the user should be able to reach the relevant source or a maintained description of it. The answer may still be wrong, but the reviewer can see what supported it and correct the right layer. This also limits false confidence. When code and documentation disagree, the system should expose the disagreement or prefer a declared authority. Quietly blending both into one smooth answer makes the conflict harder to detect. An internal agent can make restricted information easier to find, which increases the cost of a loose access model. Filtering the final answer is too late if the agent has already retrieved data the user should not see or allowed it to affect the response.
Permission checks need to travel with each source. The retrieval service should evaluate the current user's access before returning a document, table description or saved memory. Cached results need the same treatment. A cache built under a broad service account can bypass otherwise careful application permissions. There is another boundary around derived information. A user may lack access to an underlying record but still be allowed to see an approved aggregate. That rule should be explicit rather than left to the model to infer. Test permissions with paired cases. Ask the same question as two roles and confirm that each answer uses only authorised sources. Then remove access and repeat the query to check that stale embeddings, memories or caches do not preserve it. Remembering a correction can stop an agent from repeating the same mistake. It can also preserve a workaround after the underlying system changes. A saved correction should include the original issue, the corrected instruction, its source, who approved it and the scope where it applies. Some corrections are global. Others belong to one team, project or user. A correction about a temporary reporting period should expire. One about a permanent identifier format may not. Users need to inspect and edit these records. A polite conversational acknowledgement is not enough. The person correcting the answer should know whether the system saved anything, where it will apply and how to remove it.
It is also worth separating a durable correction from one session's direction. "Use the finance definition of active accounts for this analysis" may be a local instruction. "This table excludes trial accounts" sounds durable, but it still needs a source and an owner before becoming shared context. An evaluation set for an internal agent should test more than whether the model can produce a plausible answer. Include cases where similar sources conflict, a definition is stale, a user lacks permission, a remembered correction applies only to another team, or live data contradicts a cached description. Record which context the agent retrieved for each case. If an answer changes after a release, the team needs to know whether the model changed its reasoning or the retrieval layer supplied different material. OpenAI's write-up describes question and expected-query pairs for its data agent. The same idea can be adapted to other internal work: define a representative question, the sources an authorised user should reach, the sources that must be excluded and the observable checks for the result. Before connecting another source, write a small contract for it:
- What questions is this source meant to support?
- Who owns its meaning and access rules?
- How will the agent detect that it is stale?
- What citation or trace can a reviewer inspect?
- Can a user correct it, and does that correction require approval?
- What happens to saved context when access is removed?
Run those checks on one narrow task before widening retrieval. If the team cannot answer them for the first source, adding more documents will mostly give the agent more ways to be confidently inconsistent.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.