✳ the wire · analysis

OpenAI is treating Astra as its first Critical cyber model — and slowing it down, 24 hours after the leaks said next week

GPT-6 Astraconfirmedby ArtificialWatch
OpenAI is treating Astra as its first Critical cyber model — and slowing it down, 24 hours after the leaks said next week
Source imagery · verified against a primary source

Twenty-four hours ago both accounts we track said Astra was launching next week. Today OpenAI published that it is slowing the model down, and the reason outranks any launch calendar.

What OpenAI said, in its own post dated today. Internal evaluations of Astra over the past few days showed significant advances in agentic coding and cybersecurity. Those results, plus outside expert assessment, led the company to conclude — its word is "last night" — that it cannot rule out critical cyber capabilities under its Preparedness Framework.

Its own post on X puts it less tentatively: OpenAI says it is treating Astra as its first "critical" model for cybersecurity under that framework. Both statements are first-party and they are not in conflict; they are the two halves of a precautionary rule. The evaluation could not rule Critical out, so the model gets handled as if Critical is true. Anyone reporting this as "OpenAI says Astra has critical cyber capabilities" is overstating it, and anyone reporting it as a maybe is understating what OpenAI has actually committed to do about it.

Where the bar sits, because the word is doing specific work. Under that framework a model reaches the Critical cybersecurity threshold if it can find and build working zero-day exploits at every severity level across many hardened real-world systems with no human in the loop, or can plan and run end-to-end novel attacks on hardened targets given nothing but a high-level goal.

The calibration point most coverage will skip: every previous model, including GPT-5.6 Sol, was assessed at High rather than Critical. This is the first time that ceiling has been in question, and the framework has been triggered this way only once before, for biology in June 2025.

What "slowing down" concretely means. OpenAI listed the steps: isolated testing environments, restricted network and tool access, stronger model-weight protection and encryption, additional monitoring and detection, sandboxed execution. Internal activities involving Astra that do not yet meet those strengthened controls are paused. Universal monitoring now runs across every agentic application of Astra including training and evaluation, with monitors reading the model's chain of thought and able to interrupt high-risk activity. Government agencies and selected safety organisations will be brought in to test it, and third-party testing partners will be given recommended controls for higher-risk workloads.

Read that list next to yesterday's leak. The claim circulating hours before the announcement was that a checkpoint slugged "mewfour" had run inside Codex for 13 hours 11 minutes at an xhigh thinking level, with a "mewthree" tested nine days earlier for 7.5 hours. We cannot verify anyone's private telemetry and are not treating it as fact. But note what OpenAI has now said it is pausing: internal activities involving Astra that do not meet the new security bar. Long autonomous agentic runs inside a coding product are exactly that category of activity, whoever is describing them.

The parts that came from Axios rather than OpenAI. The company briefed Axios first, and a White House official told them OpenAI voluntarily informed the administration of its plans to delay the release. Axios notes this may be the first time a frontier lab has committed to slowing progress on its own model over cyber concerns — and that Anthropic, which once committed to pausing training if capabilities outran its ability to control them, rolled that commitment back in a February update to its Responsible Scaling Policy, reasoning that one developer pausing alone could leave the world less safe. Earlier this week at Black Hat, an OpenAI technical staff member described the company as consciously slowing research to improve security. All of this lands while the administration is still building a pre-release model review process whose central terms, Axios reports, were operationalised but not defined.

Also stated plainly, because it answers a question our own coverage raised: Astra was not involved in the Hugging Face exploitation. That is consistent with what OpenAI told us on July 29 about ExploitGym and again on August 4 about third-party evaluations.

What this does to the launch hunt. Nothing has changed in the catalogs — no astra, no mewfour, across 400 OpenRouter ids and 6,218 on models.dev — and our alerts for both the product name and the codename stay armed exactly as they were. What changes is the read on the date. "Next week" was always the weakest part of yesterday's leak, and it is now the least likely reading of a company that has just told the White House it intends to delay. The codename survived contact with the evidence; the calendar did not.

One note on ourselves. We published the mewfour codename yesterday as rumor and armed it, while saying the date was unverifiable. That split was the right one, and it is worth naming why: a codename is a claim a catalog can eventually confirm, and a date is a claim about a decision nobody has made yet. Today the company made the decision, and it went the other way.

Source: OpenAI's own post and X thread, read directly · Axios exclusive (briefed first) · OpenRouter + models.dev catalogs · GPT-6 Astra tracker · the bench index

sweeping every 60 seconds

Know the minute it drops — not the minute we write it up.

Claude Opus 5 went live at 16:51 UTC. The alert was in subscribers' inboxes at 16:52.

  • Free forever
  • No card
  • Unsubscribe in one click

← back to the wire