3 September 2026 / AI governance

Recheck the permission boundary when the model changes

OpenAI has classified GPT-6 Astra at its Critical cybersecurity capability threshold. A team moving to a materially more capable model should review the identity, tools and approvals around it before treating the upgrade as a drop-in replacement.

An empty white service corridor with black access doors and a red line running along the walls.
Generated image: Villar David editorial

OpenAI has classified GPT-6 Astra at its Critical cybersecurity capability threshold. A team moving to a materially more capable model should review the identity, tools and approvals around it before treating the upgrade as a drop-in replacement. The GPT-6 Astra safety overview says the model can find previously unknown security flaws and develop ways to exploit well-protected systems when it has the right tools and access. OpenAI also reports stronger resistance to prompt injection and better behaviour in browsing and workplace evaluations. Those results matter, but they don't describe the permissions in your application. A model upgrade can leave the surrounding integration unchanged while altering what the whole system can do with it. The same repository access, shell, browser session and cloud role may now support a longer or more effective chain of actions. I would treat that as a change to the system's authority, even when the API shape and tool list have not changed.

Start with the identity the agent uses. List the accounts, service roles, stored browser sessions and delegated user permissions available in each environment. For every identity, record what it can read and change. A broad development credential often survives into an agent experiment because it is convenient. That becomes harder to justify when the model is more capable of finding paths through the systems connected to it. Then separate the tools by effect. Reading a repository, proposing a patch and merging to the protected branch are different permissions. Looking up a customer record is different from changing it. Preparing an email is different from sending it. I prefer the agent to cross those boundaries explicitly so the application can require approval at the point where the consequence changes. A tool description is useful context for the model. It is not an access-control rule. The service behind the tool still needs to check the caller, tenant, object and requested operation. It should reject work outside the assigned scope even if the model asks confidently and supplies a plausible reason.

Credentials should fit the task and expire. A coding agent investigating one repository does not need a token that can administer every repository in the organisation. A support agent preparing a response does not need the same role used by a background synchronisation service. Short-lived credentials and task-specific scopes reduce what a mistaken or manipulated run can change before someone intervenes. Input trust also needs another pass. Instructions can arrive through issues, documents, web pages, emails and tool results. The model may need to read that material, but the surrounding harness should keep it separate from the authority granted by the operator. A sentence inside a retrieved page cannot approve a deployment or widen a data query. The approval has to come through the channel the application recognises for that decision.

OpenAI says Astra is more robust to prompt injection than GPT-5.6 Sol. It also reports reduced chain-of-thought monitorability in adversarial evaluations. I read those as reasons to test the behaviour we can observe directly: requested tools, arguments, changed files, external writes and the final state of the target system. Internal reasoning can help a provider investigate a model. An operator still needs evidence at the boundary where an action reaches code, data or another person. Run the old and new model against the same high-risk cases before changing production traffic. Include ambiguous instructions, a malicious document, a stale approval, a tool that returns data from the wrong account and a request that would require broader access. Check whether the system refuses the action, asks for approval or stays inside a read-only path. Keep the expected result beside the test so an improved answer cannot hide a worse permission decision.

Rollout can follow the consequence of the work. Read-only analysis can move first. Proposed changes can follow once the evidence and review path are clear. Irreversible writes, production credentials and communications to other people should wait until the new model has passed the cases that matter for those actions. For the next model upgrade, I would put the model name beside the identity and tool policy in the release record. Re-run the permission tests before switching the alias. If the model's capability has changed, the decision about what it may do should be visible too.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs