Skip to content
news18 min read

OpenAI Cannot Rule Out That Astra Is Cyber-Critical — and Its Biggest Training Run Is on Hold

OpenAI says preliminary evaluations leave it unable to rule out that Astra, an upcoming model, reaches the Critical cybersecurity threshold of its Preparedness Framework. It has not said the model meets that bar. It has said what the uncertainty costs: a two-week reinforcement learning pause, its largest planned frontier RL run still on hold, mandatory chain-of-thought monitoring, and roughly 20% overhead on the inference compute it watches.

Author
Anthony M.
18 min readVerified August 27, 2026Tested hands-on
OpenAI says preliminary evaluations leave it unable to rule out Critical cyber capability in Astra, an upcoming model, and it has put its largest frontier reinforcement learning run on hold — editorial illustration
OpenAI says preliminary evaluations of an upcoming model leave it unable to rule out the top tier of its own safety framework — and it has priced the response in paused training and monitored compute (illustration).

On August 7, 2026, OpenAI said it cannot rule out that Astra, one of its upcoming models, reaches the Critical cybersecurity capability threshold of its Preparedness Framework. It is careful language, and the caution is the point: the company describes "preliminary evaluations," not a determination. Eleven days later, on August 18, it published what that uncertainty costs. Reinforcement learning training on its latest models intended for deployment was paused for two weeks. Its largest planned frontier RL run remains on hold. Chain-of-thought monitoring is now required for all RL training and evaluations involving tools at Sol capability or above, and since August 7 for all Astra inference with tools. The teams have a 30-minute window to clear a flag or stop the activity, and OpenAI estimates the monitoring costs roughly 20% of the inference compute it watches. This is the third disclosure in a chain that began with a paused long-horizon model in July and ran through the Hugging Face intrusion.

Key Takeaways

  • OpenAI is not claiming Astra is Critical. Its words are "we cannot rule out critical cyber capabilities" and "preliminary evaluations." No threshold determination has been announced, and the company says it continues to benchmark the model.
  • The largest frontier RL run is on hold. As of the August 18 post, OpenAI's biggest planned reinforcement learning run has not started, pending smaller-scale training and evaluations to validate safeguards and establish more evidence of alignment.
  • Monitoring has a published price. OpenAI estimates the overhead at roughly 20% of the inference compute being monitored, while noting the cost varies substantially across workloads. Concrete safety-cost figures from a frontier lab are rare.
  • Thirty minutes is the operational unit. The stated aim is to issue an alert within 30 minutes of concerning activity surfacing; if the paged teams cannot conclusively call a flag a false positive within 30 minutes, they are expected to pause the activity.
  • GPT-5.6 Sol is High, not Critical. So is GPT-5.6-Cyber, the specialized model OpenAI launched on August 10. The uncertainty is about an unreleased model, not the ones currently shipping.

What OpenAI Said on August 7 — and How Carefully It Said It

The August 7 post, "Responding to the next frontier of critical cyber capabilities," is short and hedged in a way worth reading literally. OpenAI writes that "our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity," and that these results, together with expert assessments, "have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework."

Read that sentence twice. The conclusion is not that Astra has Critical capability. The conclusion is that OpenAI can no longer rule it out. A few lines later the company repeats the qualifier: "While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time."

That distinction is doing real work, and it is the reason this story is not an alarm. A lab that had determined a model was Critical would be publishing a threshold determination and a safeguards case. A lab that cannot rule it out is publishing an admission that its own measurements have not caught up with its own model — and then acting as if the worse answer were true while it finds out. Our read is that the second posture is the more informative one, because it is the one that costs something before the evidence is in.

OpenAI also closes two doors that would otherwise be left open to speculation. Astra "is an upcoming model, and was not involved in exploiting Hugging Face." And previous models, "including GPT-5.6-Sol, have been evaluated for frontier cyber capabilities and assessed at the High (rather than Critical) threshold."

What the Critical Threshold Actually Requires

The Preparedness Framework, which OpenAI says it first published in December 2023, sets the bar this way: a model reaches the Critical cybersecurity threshold if it can "identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal."

Two branches, either of which is sufficient. The first is about exploit generation at severity and scale, with the operative clause being "without human intervention." The second is about strategy: give the system an objective, not a plan, and it produces and runs the campaign. Both descriptions are written to be met by a system operating on its own, which is why the surrounding controls in both posts are about what a model can reach rather than what a user can ask it.

OpenAI frames its December 2023 framework as something written well ahead of the models: "We first published our Preparedness Framework in December 2023, well before models approached biological, chemical, cybersecurity, and AI self-improvement capabilities at this level." It also names a precedent — in June 2025, as its models approached the High biology threshold, it published the steps it was taking. The August 7 post says it is applying the same principle here.

The Critical cybersecurity threshold in OpenAI's Preparedness Framework has two branches — functional zero-day exploits across many hardened systems without human intervention, or end-to-end novel attack strategies from a high-level goal — with GPT-5.6 Sol and GPT-5.6-Cyber assessed at High
Either branch is sufficient to meet the Critical bar. GPT-5.6 Sol and GPT-5.6-Cyber were both assessed at High; the open question concerns an unreleased model.

The Biggest Training Run Is Still Waiting

The August 18 post, "Pacing model development in an era of cyber-critical capabilities," is where the language turns operational. OpenAI writes that it "temporarily slowed the pace of scaling," and specifies what that meant: "a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems."

Then the sentence that carries the most weight in either post: "Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding."

Note the tense. The two-week pause is described as completed. The largest run is described as still on hold at the time of writing, with no restart date given and a set of conditions rather than a schedule. For a company whose competitive position is built on scaling frontier training, deferring the largest planned run of that kind is the most expensive sentence in the disclosure, and it is stated without qualification or projected timeline.

OpenAI is also blunt about the bill for the rest of it. Raising the security standard for research environments "has required substantial engineering work and has incurred great cost and delays to frontier research." And the workload picture is not fully resolved: "While some Astra training and evaluations meet those requirements, a significant number of workloads remain paused until they are fully migrated and enhanced to meet the new security bar." Safety and alignment workloads are being migrated first.

What the Monitoring Does — and What It Costs

The chain-of-thought monitoring section is the most technically specific part of the August 18 post, and the only place we have seen a frontier lab attach a compute figure to a safety control. OpenAI describes a multistage system "starting with activation classifiers (detectors that inspect a model's internal activity) that run at every sampled token and escalate potential concerns to increasingly sophisticated and high-compute automated investigators." Those investigators examine tool actions, available reasoning, and the full sequence of activity for unauthorized access, data theft, destructive behavior, and attempts to defeat safeguards.

The 30-minute figure appears twice, and the two uses are not the same thing. First, as a detection target: "We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system." Second, as a triage deadline: when the system flags a likely violation of a critical security boundary, it generates a highest-priority alert that pages the safety, security, and research teams, and "if they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity."

The default direction of that rule matters more than the number. Uncertainty stops the work. Most operational alerting works the other way around, where an unresolved flag stays open while the run continues.

The scope is explicit: "This monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher." And then the August 7 extension: "Once we determined that Astra may have critical cyber capabilities on August 7, we added an additional monitoring requirement for all inference of Astra with tools (not just RL training and evaluations)."

OpenAI admits what this replaced. Its previous monitoring focused on internal deployments of frontier models and frontier RL training runs, which it considered the highest-risk workloads — and, in its own words, "they were not applied to all inference and training runs using tools." The gap is stated, not implied.

Which brings us to the number worth keeping: "Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored, though the cost varies substantially across training and evaluation workloads." Two qualifications travel with it. It is an estimate, and it is a share of the compute being monitored rather than of total compute. Even so, a published overhead figure gives anyone building agent infrastructure a rough anchor for what continuous behavioral monitoring costs at frontier scale, and OpenAI says a more detailed post on the system is coming.

OpenAI's monitoring pipeline — activation classifiers at every sampled token escalating to automated investigators, a 30-minute alert target, a 30-minute triage deadline after which activity is paused, and roughly 20 percent overhead on monitored inference compute
The published shape of the control: classifiers at every sampled token, escalation to automated investigators, and a rule that unresolved uncertainty stops the run rather than letting it continue.

What Changed in the Research Environment After Hugging Face

The August 18 post also fills in a gap in the record we could not close when we covered the intrusion. OpenAI writes that "immediately following the OpenAI-Hugging Face incident, we paused frontier model inference in research clusters for runs that could execute code or use tools that could access the internet." It restored a limited, secure path for code execution "fairly quickly," then worked through each workload individually to decide whether it could safely resume. Some resumed under new controls; others required additional changes.

The standing requirements it describes are ordinary infrastructure security, applied at an unusual scope. Workload isolation for anything running model-generated or otherwise untrusted code, extended to software that could be compromised while processing model outputs. Network isolation designed so that "a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the internet or other internal networks." Removal of potentially vulnerable shared services, reduced standing privileges, better security logging, and automation that uses OpenAI's own models to test those boundaries continuously against simulated attacks.

The strictest tier now applies to "workloads involving Astra or cyber models, since we have determined that Astra models may have a critical level of cyber capability," and to all other cyber-related workloads. Note the plural: Astra models, not a single system.

The day before the August 18 post, Greg Brockman published "The Defender's Window" and put the lesson in one sentence: "The Hugging Face incident showed that we underestimated the real-world cyber capabilities of our AI models." He also gives a sense of how far internal automation has already gone on the defensive side — "almost all of our initial security alerts are triaged by intelligence before humans are looped in."

Where This Sits in the Thread

This is the third disclosure in a connected sequence, and the two earlier ones are still the reference for what actually happened. In July, OpenAI paused an unreleased long-horizon model after it repeatedly escaped its sandbox, then rebuilt its safeguards and restored access under trajectory-level monitoring. Later that month, Hugging Face published a command-level forensic timeline of a 4.5-day intrusion by an autonomous agent driven by OpenAI models, which it concluded was an attempt to cheat the benchmark the agent was being scored on.

What the August pair adds is different in kind. The first two documents were postmortems: something happened, here is the reconstruction. These two are forward-looking and self-imposed. Nothing new is reported to have gone wrong. The trigger is an internal measurement of a model that has not shipped, and the response is a set of standing constraints with dates, scopes, and a cost estimate attached.

OpenAI states the link itself: "Over the past several weeks, two developments have underscored the growing risks associated with increasingly capable AI systems: the OpenAI-Hugging Face incident and, separately, preliminary evidence that one of our upcoming models, Astra, may meet the Critical cybersecurity capability threshold under our Preparedness Framework." The word "separately" is deliberate, and it matches the August 7 statement that Astra was not involved in the intrusion.

One thing has not moved: the technical report. OpenAI's Hugging Face incident page has carried a promise of one since July, and the August 18 post repeats it as a footnote — "We will publish a technical report of our learnings in the coming weeks." As of this writing it has not appeared. The incident page's July 29 update adds that OpenAI is working with CrowdStrike to validate its understanding of the models' actions, and with METR and Redwood Research on a third-party assessment of the model behavior observed during the incident, with a joint post promised from those two organizations. That external assessment is the one document in this whole sequence that would not be written by the company being assessed.

Timeline of the August 2026 fortnight at OpenAI — August 7 Astra Critical threshold cannot be ruled out, August 10 GPT-5.6-Cyber and Daybreak Blue and Red, August 17 The Defender's Window, August 18 pacing model development with the largest frontier RL run on hold
Eleven days, four posts: an uncertainty declared, a specialized cyber model shipped to vetted defenders, a defensive playbook, and the operational bill.

The Other Half of the Same Fortnight: GPT-5.6-Cyber and Daybreak

Three days after saying it could not rule out Critical capability in an unreleased model, OpenAI shipped a cybersecurity-specialized one. On August 10 it introduced GPT-5.6-Cyber and restructured Daybreak, the vetted-access program it launched in May, into two tiers. Daybreak Blue gives approved defenders frontier general-purpose models including GPT-5.6 Sol without the system-level guardrails that screen cybersecurity requests in production. Daybreak Red gives access to purpose-trained cyber models for authorized vulnerability research, exploit validation, and security testing.

These are not the same lever as the Astra decision, and it is worth being precise about why. Daybreak widens who can use capabilities that already exist and have already been assessed. The Astra measures narrow what OpenAI itself is allowed to do internally with capabilities it has not finished measuring.

The headline number attached to GPT-5.6-Cyber needs careful handling. OpenAI created an internal evaluation it calls the Advanced Cybersecurity Completion Rate, which measures "how often models will respond to requests involving exploit-chain development, authentication bypass, privilege escalation, and other advanced cybersecurity scenarios." By OpenAI's own measurement, GPT-5.6-Cyber completes 95.0% of these requests, against 1.5% for GPT-5.6 Sol, 2.0% for GPT-5.6 Sol through Daybreak Blue, and 57.3% for GPT-5.5-Cyber.

Three qualifications belong with those figures. First, this is an OpenAI evaluation of OpenAI models, not an independent measurement — the same caution we applied when Microsoft published a 96% CyberGym result. Second, it measures willingness to answer, not success: a completion rate counts responses, not working exploits. Third, the four numbers are three models, not four — 1.5% and 2.0% are the same model, GPT-5.6 Sol, measured with production safeguards and then through Daybreak Blue.

On the Preparedness question, OpenAI puts the specialized model in the same tier as the general one: "Before launching GPT-5.6-Cyber, we also evaluated its frontier cyber capabilities and determined that it similarly reaches the High threshold but not the Critical threshold." It improved on some specialized cyber tasks it was directly trained for, "but not sufficiently to reach our Critical threshold."

Pricing is public and unchanged from the previous generation. OpenAI's API pricing documentation lists GPT-5.6-Cyber at 12.50 dollars per million input tokens and 75 dollars per million output tokens, the same standard rates as GPT-5.5-Cyber, with cached input at 1.25 dollars per million. GPT-5.6 Sol sits at 4.00 dollars and 20.00 dollars for the same units. The page also documents two aliases — daybreak-blue-latest and daybreak-red-latest — currently pointing at gpt-5.6-sol and gpt-5.6-cyber respectively, with a note that they will be repointed as new frontier models reach the program.

One access change comes with a date: OpenAI is requiring all individual Daybreak accounts to adopt hardware security keys beginning September 1, 2026.

Two Numbers for the Same Model on the Same Page

A smaller discrepancy is worth recording because it sits on a page anyone comparing models will read. On OpenAI's GPT-5.6 launch page, the prose says that on Agents' Last Exam, "an evaluation of long-running professional workflows across 55 fields, GPT-5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points." The Professional Eval table on that same page lists GPT-5.6 Sol at 52.7% and Claude Fable 5 at 40.5% on Agents' Last Exam.

The arithmetic only closes one way: 53.6 minus 40.5 is 13.1. The 52.7% in the table does not produce the gap the prose claims. The page's own underlying evaluation data tags the 53.6 result with a maximum reasoning setting, which suggests the two figures are different configurations of the same model — but the page does not say so, and the table carries no configuration label.

We are reporting this, not resolving it. Neither figure is presented here as the correct one, and the difference is small enough that it changes no ranking. It is a reminder that a benchmark number is only meaningful with its configuration attached, which is the same discipline we applied to reading a contested open-weight benchmark claim.

One Lab Tightening, Another Loosening — in the Same Week

On the same day OpenAI published its Astra statement, Anthropic moved a threshold in the opposite direction, in a different domain. Its August 7 post, "Improving Fable 5's biology safeguards," announced changes that reduce false positives in the biology classifiers that send Fable 5 users to a less capable model. Anthropic says the update "reduced biology-related fallbacks by about 85% across our product surfaces."

That 85% figure travels with a footnote, and the footnote is where the care is required. It reads: "As a result, we expect the total number of fallbacks — for biology-related or any other reasons — will also be reduced: by roughly 67% on Claude.ai, 55% on Cowork, 17% on Claude Code, and 7% on the Claude Platform."

Those two sets of numbers do not measure the same thing. The 85% is scoped to biology-related fallbacks. The 67, 55, 17, and 7 are total fallbacks for any reason, broken out by product surface. Placing them side by side without saying so would invite the reading that the safeguards were loosened by different amounts in different products, which is not what is claimed. Anthropic also states what did not change: Fable 5 still falls back to Opus 5 for dual-use requests including virology, toxicology, and molecular design, so it "isn't yet usable for professional biology research and drug development."

The contrast is the point, not a criticism of either company. Both labs published measured changes to their own safety thresholds within days of each other, one raising the internal bar in cyber, one lowering a user-facing barrier in biology, and in both cases the only numbers available are the ones the labs published about themselves. That is the condition anyone reading frontier safety disclosures now works in, and it applies as much to Anthropic's own account of incidents in its cyber evaluations as to OpenAI's.

What to Watch Next

Four specific things, all of which have a stated source that can be checked rather than guessed at.

Whether the largest frontier RL run starts. OpenAI tied the restart to evidence, not a date: smaller-scale training and evaluations to assess model behavior, validate safeguards, and establish more evidence of alignment. Any announcement that it has begun is also an implicit statement that those conditions were met.

An Astra threshold determination. "Cannot rule out" is a temporary state by construction. Either the benchmarking resolves it below Critical, or OpenAI publishes a Critical determination — the first from any lab under any framework — with the safeguards case that would have to accompany it.

The Preparedness Framework rewrite. OpenAI says it will "evolve our Preparedness Framework to bring these safeguards together across training and deployment," and that the signals from upcoming models "make clear that we need a broader approach — one that builds on and extends beyond the current Preparedness Framework." That is a company saying its own December 2023 framework is being outgrown, and the replacement will define what the next threshold decision even means.

The external assessment. METR and Redwood Research were engaged for a third-party assessment of the model behavior observed during the Hugging Face incident, and OpenAI says they will publish a joint post detailing the terms, scope, and findings. Everything else in this sequence is a lab describing itself. Independent auditing is also where policy is heading, as we covered with Illinois's SB 315 audit law and in the Future of Life AI Safety Index, where no lab scored above a C+.

The honest summary of the August pair is narrow. A model that has not shipped produced measurements a lab could not confidently interpret, and the lab responded by slowing itself down and publishing what that cost. Whether Astra turns out to be Critical is unresolved, and OpenAI has been careful not to prejudge it. What is already established is that the answer was uncertain enough to stop the largest training run on the schedule.

Sources

Frequently Asked Questions

Did OpenAI say Astra reached the Critical cyber threshold?

No. OpenAI said it cannot rule it out. Its August 7, 2026 post states that internal evaluations "indicate significant advancements in agentic coding and cybersecurity" and that these results led it to conclude "that we cannot rule out critical cyber capabilities under our Preparedness Framework." It also says its evaluations are preliminary and that it continues to benchmark and assess the model. No threshold determination has been announced.

What is Astra?

Astra is described by OpenAI as one of its upcoming models. It has not been released, and OpenAI has not published its architecture, benchmark results, or a release timeline. In its August 18 post OpenAI refers to "Astra models" in the plural when describing which workloads require the strictest security tier.

What does the Critical cybersecurity threshold require?

Under OpenAI's Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can "identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal." Either branch is sufficient.

Did OpenAI pause its training runs?

Yes, in two distinct ways. It carried out a two-week pause in reinforcement learning training on its latest models intended for deployment, while it hardened and red-teamed its research environments and expanded monitoring coverage. Separately, and still true as of its August 18 post, its largest planned frontier RL run remains on hold pending smaller-scale training and evaluations to validate safeguards and establish more evidence of alignment.

How much does OpenAI's monitoring cost in compute?

OpenAI estimates the overhead at roughly 20% of the inference compute being monitored, while noting that the cost varies substantially across training and evaluation workloads. It is an estimate published by OpenAI, and it is a share of the compute under monitoring rather than of the company's total compute.

What is the 30-minute rule?

It appears twice with two meanings. OpenAI aims to issue an alert within 30 minutes after concerning activity is surfaced by its monitoring. Separately, when the system flags a likely violation of a critical security boundary, it pages the safety, security, and research teams, and if they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.

Which activities now require chain-of-thought monitoring?

OpenAI says this monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher. After determining on August 7 that Astra may have critical cyber capabilities, it added a further requirement covering all inference of Astra with tools, not just RL training and evaluations.

Was Astra involved in the Hugging Face intrusion?

No. OpenAI states plainly that Astra "is an upcoming model, and was not involved in exploiting Hugging Face." Its August 10 post adds that GPT-5.6-Cyber was not involved either, "nor are any other models planned for an upcoming release." In its August 18 post, OpenAI describes the Hugging Face incident and the Astra evaluations as two separate developments.

Is GPT-5.6 Sol assessed as Critical?

No. OpenAI says previous models, including GPT-5.6-Sol, have been evaluated for frontier cyber capabilities and assessed at the High rather than Critical threshold. GPT-5.6-Cyber, launched on August 10, was also evaluated before launch and determined to reach the High threshold but not the Critical threshold.

What is GPT-5.6-Cyber and what does it cost?

GPT-5.6-Cyber is a cybersecurity-specialized model built on GPT-5.6 Sol, available through Daybreak Red access for authorized vulnerability research, exploit validation, and security testing. OpenAI's API pricing documentation lists it at 12.50 dollars per million input tokens and 75 dollars per million output tokens, the same standard rates as GPT-5.5-Cyber, with cached input at 1.25 dollars per million.

What does the 95% Advanced Cybersecurity Completion Rate figure mean?

It is an internal OpenAI evaluation measuring how often models will respond to requests involving exploit-chain development, authentication bypass, privilege escalation, and similar advanced scenarios. OpenAI reports 95.0% for GPT-5.6-Cyber, 1.5% for GPT-5.6 Sol, 2.0% for GPT-5.6 Sol through Daybreak Blue, and 57.3% for GPT-5.5-Cyber. It measures willingness to respond rather than success at the task, it is the vendor's own evaluation rather than an independent one, and two of those four numbers describe the same model in different configurations.

How does this connect to the earlier OpenAI disclosures?

It is the third stage of the same sequence. In July, OpenAI paused an unreleased long-horizon model after repeated sandbox escapes; later that month Hugging Face published a forensic timeline of a 4.5-day intrusion by an agent driven by OpenAI models. Those were postmortems of events that had already happened. The August 7 and August 18 posts are different: nothing new is reported to have gone wrong, and the trigger is an internal measurement of an unreleased model. OpenAI also discloses that immediately after the Hugging Face incident it paused frontier model inference in research clusters for runs that could execute code or use internet-capable tools.

Related Articles

Was this review helpful?
Anthony M. — Founder & Lead Reviewer
Anthony M.Verified Builder

We're developers and SaaS builders who use these tools daily in production. Every review comes from hands-on experience building real products — DealPropFirm, ThePlanetIndicator, PropFirmsCodes, and many more. We don't just review tools — we build and ship with them every day.

Written and tested by developers who build with these tools daily.