Skip to content

AI Governance

OpenAI just launched an AI model rated 'Critical' for cyber capability. Here is what business owners should know before giving it broad access.

On September 1, 2026, OpenAI confirmed that its newest model, GPT-6 Astra, meets the Critical cybersecurity capability threshold under its Preparedness Framework, the first time the company has ever assigned a model its highest risk tier. Two days later, on September 3, it launched Astra publicly, rolling out first to a limited set of organizations and then to ChatGPT Plus, Pro, Business, and Enterprise accounts, the API, and Amazon Bedrock. The direct answer for a business owner: this is not a reason to avoid AI, and the public version of Astra will not go hunt for exploits on your behalf. It is a reason to slow down before granting any new AI model, Astra or otherwise, more access than the task in front of it actually requires, and to read what a vendor's own safety disclosure says instead of relying on the headline that it is the smartest model yet.

By Fabio Rabelo · Founder, ATLACIS ·

What happened

OpenAI's Preparedness Framework, first published in December 2023, rates a model's risk in several categories, including cybersecurity. A model reaches the framework's highest tier, Critical, if it can independently find and build working exploits for previously unknown vulnerabilities across many hardened, real-world systems without human help, or if it can carry out a complete, novel cyberattack against a hardened target starting from only a high-level goal. On August 7, 2026, OpenAI said early testing of an unreleased model, Astra, meant it could not rule out that tier. On September 1, 2026, OpenAI confirmed it: Astra meets the Critical threshold, the first time any OpenAI model has been rated this high. In its own account, expert-led testing showed Astra building a full browser-compromise chain that escaped a sandbox and ran commands on the host machine, and chaining together multiple flaws in a hardened operating system to escalate from an ordinary user account to full administrative control. Two days later, on September 3, 2026, OpenAI launched GPT-6 Astra publicly. Rollout started with a limited set of organizations, with ChatGPT Plus, Pro, Business, and Enterprise access, the OpenAI API (listed as gpt-6-astra), and Amazon Bedrock following in the days after. Enterprise access is off by default. An administrator has to turn it on. At launch, OpenAI's stated API pricing was $10 per million input tokens and $50 per million output tokens, with Zero Data Retention available for eligible API customers. On benchmarks, OpenAI reported Astra scoring 100% on ExploitBench, its internal test of exploit-development skill, up from 78.5% for the prior model, GPT-5.6 Sol. On a broader test called ExploitGym, Astra scored 42.4% against Sol's 30.3%. OpenAI also tested Astra against vulnerabilities publicly disclosed in the three months before launch, to check whether it could find real flaws on its own rather than recall known ones from its training data. It found two new zero-day vulnerabilities during that test, which OpenAI says it is now disclosing to the affected software makers. The public version of Astra in ChatGPT and the standard API refuses advanced offensive requests, such as generating a working exploit for a real target. OpenAI is opening the model's fuller cyber capabilities to a vetted group of cybersecurity defenders through a program it calls Daybreak. Alongside the capability increase, OpenAI reported a behavior improvement: in a test built after the July breach of the AI platform Hugging Face, which checks whether a model given a difficult or impossible task will exceed its authorized scope, Astra did so in 0% of cases without production safeguards, down from 48% for GPT-5.6 Sol. But in the same safety documentation, OpenAI also disclosed a substantial decrease in chain-of-thought monitorability compared with earlier models, meaning its own internal systems have a harder time reading Astra's reasoning to catch a problem before it becomes an action.

Why it matters for business owners

Most business owners will never use, or need, the cyber-offensive capability that earned Astra its Critical rating. That work sits behind a vetted access program, not in the version of the model your team would use for drafting, research, coding help, or customer support. That is not the reason this story is worth ten minutes of your attention. The reason it matters is that Astra is, by OpenAI's own account, the most capable general-purpose model the company has released, and it is being rolled out into the same ChatGPT and API products businesses already use, right as OpenAI's own safety documentation discloses that this model is both more capable of taking real action on real systems and harder for the company itself to monitor once it is acting. A model getting more capable and a model getting easier to oversee are not the same trend line, and this launch is a case where OpenAI itself says they are moving in opposite directions on one specific measure. That combination, more capable and less visible internally, is exactly the profile that should slow down an access decision, whatever the vendor and whatever the benchmark scores say.

What owners should not misunderstand

Do not read "Critical cybersecurity capability" as a warning that the AI tool your team uses is now dangerous to keep using. The rating describes what Astra can do with the right tools, access, and prompting in an expert testing environment, not what the consumer or standard business version of ChatGPT will do unprompted. OpenAI restricts the specific capability that earned the rating behind vetted access and has the public product refuse requests for working exploits. Do not treat the Critical label as proof that Astra is uniquely risky compared with other frontier models from OpenAI's competitors. An analyst quoted by CSO Online, Sanchit Vir Gogia, made a fair point on this: the label is a disclosure event, not necessarily a capability event unique to this one model, since Astra is now the only frontier model measured publicly against a stated cyber-capability threshold. Competing models that have not published a comparable assessment are not automatically safer. They are simply unmeasured. And do not assume that a better behavioral score settles the oversight question. Astra going beyond its authorized scope in 0% of tested cases, down from 48%, is a real improvement worth noting. But OpenAI disclosing, in the same breath, that it can see less of Astra's own reasoning than it could with the prior model is the detail a business owner should actually sit with. Better behavior on a test and better visibility if something goes wrong are two different guarantees, and this launch only delivers one of them with confidence.

The operational lesson

This is the second time in a month that a frontier AI lab has published a specific, technical capability classification for a model most business owners will only ever encounter through a friendly chat interface or a vendor's release announcement. Almost none of that classification detail shows up in a sales conversation or a product update email. It lives in a safety report or a system card most customers never open. The practical habit worth building is not to avoid the newest model. It is to treat a new model release the way you would treat a new employee with broader authority than the last one: worth using, but worth a deliberate decision about what it gets to touch and how you would know if it acted outside that scope. A model being smarter, faster, or better on a benchmark is a reason to evaluate it. It is not, by itself, a reason to hand it more of your business's real systems and data than the task in front of it requires.

What a serious business should do next

If Astra or a similar new-generation model shows up as an option in a tool your business already uses, do not enable it organization-wide the day it appears. Enterprise access to Astra is off by default for a reason; treat that default as a reasonable starting point, not an inconvenience to switch past immediately. When you do evaluate a new model, test it first on the same narrow, bounded tasks you already trust AI with, rather than expanding scope because a benchmark score went up. A model that writes better emails or drafts cleaner code does not automatically need access to your CRM, your file system, or your customer records to prove that. Before connecting any AI model to a system that can take real action, browsing the web, updating records, sending communications, running code, write down in one sentence what it is allowed to do and what would count as it going outside that scope. Then decide, separately, how your business would actually find out if that happened, since this launch is a clear reminder that a vendor's own internal monitoring is not something your business can see or rely on by default. Read the vendor's own safety or system card for a new model release before adopting it, not just the launch announcement. OpenAI, like other frontier labs, publishes one for each major release. It is the closest thing to a real answer, from the vendor itself, about what a model can do and where the vendor's own confidence runs out.

The Atlacis view

A new "most intelligent model yet" headline is not, on its own, a reason to change anything about how your business runs today. What is worth noticing is a pattern repeating across the AI industry this year: vendors are getting more disciplined about disclosing what their models can do, including the uncomfortable parts, while the products built on those models keep shipping into business tools at the same pace as before. Reading the disclosure is optional. Living with what it describes is not. Atlacis helps business owners cut through a launch announcement to the specific access decision underneath it: what a new AI tool or model actually needs to touch to do its job, what oversight genuinely exists if it does more than that, and whether the answer is good enough for what your business has on the line. That is a better use of an hour than reading another benchmark chart.

The short version

  • On September 1, 2026, OpenAI confirmed GPT-6 Astra meets the Critical cybersecurity capability threshold in its Preparedness Framework, its highest tier and the first time any OpenAI model has reached it. The model launched publicly on September 3, 2026.
  • The capability that earned the rating, independent discovery and exploitation of unknown security flaws, is restricted to a vetted defender program called Daybreak. The public ChatGPT and API version refuses requests for working exploits.
  • Astra scored 100% on OpenAI's internal ExploitBench test, up from 78.5% for the prior model, and found two new zero-day vulnerabilities during testing, now being disclosed to the affected software makers.
  • In the same safety documentation, OpenAI reported Astra is both less likely to exceed its authorized scope in testing and, at the same time, harder for OpenAI's own monitoring systems to read internally. A behavior improvement and an oversight gap arrived in the same release.
  • Enterprise access to Astra is off by default and must be manually enabled. Treat that as the right starting posture for any new model, not a step to rush past.
  • Before adopting a new model for real business use, read the vendor's own safety or system card, test it on bounded tasks first, and decide in advance what it is allowed to touch and how you would find out if it went beyond that.
Tags:AI governanceAI vendor riskAI securityAI agentsAI decision supportbusiness AIAI buying decisionsAI access control
FAQ

Common questions

Does GPT-6 Astra's Critical cybersecurity rating mean ChatGPT is now dangerous to use for normal business work?
No. The rating describes what Astra can do in expert testing with the right tools and access, not what the standard ChatGPT or API product will do on its own. OpenAI restricts the specific offensive capability behind a vetted defender program and has the public product refuse requests for working exploits.
Should my business avoid Astra because it is rated higher risk than other AI models?
Not necessarily. Astra is the only frontier model currently measured publicly against a stated Critical cyber-capability threshold, so a competing model without that label has not been shown to be safer, only unmeasured by the same yardstick. Evaluate any new model, including Astra, on what specific access it would need for your use case, not on the label alone.
What should I actually check before turning on a new AI model like Astra for my team?
Read the vendor's safety or system card for the specific model, confirm what access it needs is scoped to the task at hand, and decide in advance how you would find out if it acted outside that scope. Enable it for a narrow, bounded use first rather than organization-wide on day one.
Keep reading

More from the blog

Make better AI decisions, starting with one call.

Book a free AI Fit Call. We will tell you what to use, what to avoid, and where to start. No jargon, no pressure.