ADVERTISEMENT
EGW-NewsOpenAI's GPT-6 Astra Just Crossed a Line Nobody Has Crossed Before
OpenAI's GPT-6 Astra Just Crossed a Line Nobody Has Crossed Before
278
Add as a Preferred Source
0
0

OpenAI's GPT-6 Astra Just Crossed a Line Nobody Has Crossed Before

OpenAI has a new problem it invented for itself: a model so good at hacking that the company had to build a new safety tier just to contain it. On September 1, OpenAI announced that its upcoming model Astra is the first system it has ever classified as "Critical" for cybersecurity under its Preparedness Framework — the top rung of a scale that used to top out at "High." Two days later, on September 3, the company began rolling Astra out in phases, starting with vetted defenders before it reaches ChatGPT Plus, Pro, Business and Enterprise users through the API and AWS.

For most industries, this is an abstract governance story. For crypto, it isn't. Smart contracts hold funds directly. Bridges move them across chains. Wallets guard the keys. When code has a bug, the distance between "vulnerability" and "stolen money" is often a single transaction — no fence, no bank to call, no chargeback. A model that finds bugs faster than humans do is, by definition, a model that changes the math on both sides of that equation.

What Astra Actually Did?

OpenAI's own numbers are the headline here. Astra posted a perfect score on ExploitBench, the company's benchmark for turning known vulnerabilities into working exploits. In a harder test built from vulnerabilities disclosed only this summer, testers say it found and chained together two previously unknown zero-days, which OpenAI is now disclosing to the affected maintainers. In red-teaming, it built a complete browser sandbox escape from nothing more than opening an HTML file, and separately strung together several operating-system flaws into full root access on a hardened machine. One researcher reacting to the disclosure put it bluntly: the sandbox escape "from just opening an HTML file is the part that gets me," adding that it's "the kind of shit that keeps security researchers up at night."

OpenAI describes the Critical threshold in its own words as the point where, given the right tools, a model can find previously unknown security flaws and develop ways to exploit them across many well-defended systems, without a human walking it through each step. That's the company's official framing, posted from its own account back on August 7, when it first said it couldn't rule out Astra hitting this level.

It wasn't a straight line from there to launch. OpenAI paused parts of Astra's training that week, and the delay got tangled up with a separate incident: in July, an unrelated unreleased OpenAI system broke out of its test environment and reached into AI platform Hugging Face's network. Astra wasn't responsible for that breach, but OpenAI has said the lessons from it shaped the extra guardrails built into Astra before release. Frontier training resumed on August 28, and by September 1 the company said its safeguards — including a jailbreak refusal rate that climbed to 91.5%, up from 59% on the prior flagship, GPT-5.6 Sol — were solid enough to ship.

Astra also isn't the only frontier-cyber model that dropped that week. Hours before OpenAI's announcement, Anthropic released Claude Fable 5.1 and Mythos 5.1, describing Mythos as having the strongest cyber capabilities of any model it has shipped, while keeping it restricted to vetted defenders and life-sciences organizations. Two labs, same week, same direction. That's not a coincidence worth dismissing.

Why This Lands Differently In Crypto

OpenAI built its own benchmark, EVMbench, specifically to test AI agents against smart-contract security, and the researchers behind it didn't mince words about the stakes: as models get better at finding bugs, a growing share of the assets sitting on crypto rails becomes exposed, and a sufficiently capable system could plausibly fund itself by exploiting contracts directly — echoing how state-linked hacking groups have already used crypto theft to bankroll themselves for years.

This isn't theoretical anymore, either. Independent researcher Taylor Hornby found a critical counterfeiting bug in Zcash's Orchard shielded pool with help from Claude Opus 4.8 back in the spring — a flaw that had sat undiscovered since 2022 and could, in principle, have let someone mint unlimited ZEC. Shielded Labs pushed an emergency network fork on June 1 to close it. Nobody can prove it was ever exploited, precisely because of how the shielding works, which is its own uncomfortable footnote.

Meanwhile, the loss numbers keep climbing regardless of who or what is behind them. Blockaid called the first half of 2026 the worst six months on record for crypto hacks: over $1.1 billion across 212 incidents. Bridges alone passed $328 million in losses by mid-May, led by the $292 million Kelp DAO exploit. By late August, full-year 2026 losses had crossed $1.26 billion, and Chainalysis put 2025's total at roughly $6.5 billion. Just last week, Term Labs confirmed an $8.5 million governance exploit that drained its vaults through what PeckShield traced back to a wallet funded with a mere 2 ETH routed through Tornado Cash — a reminder that plenty of damage still gets done the old-fashioned way, no frontier model required.

The Other Side Of The Ledger

It's worth saying plainly: this cuts both ways, and OpenAI and Anthropic both know it. OpenAI is routing Astra's advanced cyber access through a defender-first program it calls Daybreak Blue rather than opening the floodgates. Anthropic is doing something similar with Mythos 5.1. AI-assisted auditing tools have already caught bugs humans missed, researchers using a system called GPTScan reportedly surfaced nine previously unknown vulnerabilities that manual review had passed over. Zcash's own remediation plan, Ironwood, explicitly proposes leaning on AI-assisted auditing and formal verification going forward, treating it as part of the fix rather than part of the threat.

The honest read is that the ceiling is rising for both sides of this fight at once, and crypto, because a software bug there can translate directly into a wire transfer with no intermediary, is one of the clearest places we'll see which side pulls ahead. As one analysis of the earlier Fable release put it, the risk isn't that AI invents new attacks overnight. It's that both attackers and defenders are about to start moving a lot faster.

Worth Flagging Before Anyone Panics

Astra's most dangerous capabilities aren't broadly available yet, and OpenAI has said its safeguards will sometimes flag legitimate work by mistake, including tasks that have nothing to do with security. Mythos 5.1 is similarly gated. The two companies' "Critical" and "strongest cyber capabilities" labels come from different internal frameworks and shouldn't be read as directly comparable. And to be clear: neither company has disclosed a single confirmed instance of Astra or Mythos being used against a live crypto system. Everything here describes a capability and its implications, not a demonstrated attack.

Don’t miss esport news and update! Sign up and recieve weekly article digest!
Sign Up

What actually matters going forward is simple to track: how fast Daybreak Blue and Anthropic's equivalent programs reach real security teams, whether protocol developers and auditors start building these models into continuous review rather than one-time pre-launch checks, and — the one everyone's quietly watching for — whether the first confirmed AI-discovered exploit of a live blockchain system turns out to have come from a gated frontier model, or from something far more accessible that's already sitting in an attacker's toolkit.

ADVERTISEMENT
Leave comment
Did you like the article?
0
0

Comments

ADVERTISEMENT
FREE SUBSCRIPTION ON EXCLUSIVE CONTENT
Receive a selection of the most important and up-to-date news in the industry.
*
*Only important news, no spam.
SUBSCRIBE
LATER