Scale · founder · 8 min read

GPT-6 Astra is here: why the biggest model launch of the year probably isn't for you

OpenAI shipped GPT-6 Astra on September 3. It's a computer-use model, not a coding upgrade, and it trails Claude Fable 5.1 on OpenAI's own index.

OpenAI released GPT-6 Astra on September 3, 2026. It is the company’s flagship model, it came with the most aggressive framing OpenAI has used for any release, and it arrived in a week where three other frontier labs also shipped.

We’re writing this a week late, and that’s deliberate. The launch-day version of this article would have been a benchmark table. A week on, the useful version is different: Astra is not a coding upgrade, and for most people reading this site, the correct action is nothing. Here’s the reasoning, because the reasoning is more useful than the verdict.

What Astra actually is

Astra operates software through screens and controls rather than through APIs. That’s the whole pitch. OpenAI’s demos show it filling in web forms, updating CRM records, editing spreadsheets, running research, and driving engineering software like KiCad and FreeCAD — clicking through real interfaces the way a person does, instead of needing an integration built for every tool.

The numbers OpenAI published for this are genuinely strong. On OSWorld 2.0 (a benchmark for operating a desktop), Astra scores 72.6% against 65.7% for GPT-5.6 Sol and 70.2% for Claude Opus 5. More interesting than the score is the clock: Astra takes about 40 minutes per task where Sol takes about 75. Roughly half the time for a better result.

If you’ve ever paid someone to copy data between two systems that don’t talk to each other, you can see why this matters. That’s the category Astra is aimed at.

The part the headlines skipped

Here is the table that should shape your read. Every number below is OpenAI’s own, from OpenAI’s own launch page.

BenchmarkGPT-6 AstraBest Claude
OSWorld 2.0 (operating a desktop)72.6%Opus 5: 70.2%
Terminal-Bench 4.057.9%Fable 5.1: 55.8%
DeepSWE v1.1 (fixing real bugs)74.1%Opus 5: 73.7%
Artificial Analysis Intelligence Index v4.1.161.2Fable 5.1: 65.7
Humanity’s Last Exam, with tools57.2%Fable 5.1: 65.0%
Price per million tokens (in / out)$10 / $50Fable 5.1: $10 / $50

Read the bottom half of that table again.

On DeepSWE v1.1 — the coding benchmark most teams actually watch — Astra scores 74.1% against 73.7% for Claude Opus 5 and 73.8% for Gemini 3.8 Flash. That is a three-way tie dressed up as a win. On the independent Artificial Analysis index, Astra sits behind Claude Fable 5.1, 61.2 to 65.7. On Humanity’s Last Exam with tools, it’s behind again, 57.2% to 65.0%.

And the price is identical to Fable 5.1. Same $10 in, $50 out.

So the honest summary is: Astra is clearly ahead on operating a computer, on cybersecurity, and on hard mathematics. On general reasoning and everyday coding it is level with the field, at the same price, and behind on the one independent index OpenAI chose to publish.

One more caveat worth internalising. OpenAI ran most of the competing models itself, and its own footnotes concede that the Claude results on some benchmarks use settings that differ from Anthropic’s reported runs. Treat every gap in that table as directional until someone outside OpenAI reproduces it.

The thing nobody is talking about

Astra is the first OpenAI model ever rated Critical for cybersecurity under the company’s Preparedness Framework. It scores 100% on ExploitBench — turning known flaws into working exploits — against 78.5% for GPT-5.6 Sol.

That rating is why you probably can’t use it yet. Access started with a small set of enterprise customers in OpenAI’s Daybreak programme, with paid ChatGPT plans and the API following “over the coming days.” Enterprise administrators have to switch it on, and it’s off by default. As of writing there is still no general availability date.

This is new. A frontier lab shipped its flagship and then deliberately throttled who can touch it, on security grounds, and said so out loud. Whatever you think of OpenAI, that’s a meaningful precedent — and it’s a preview of what “model launch” is going to mean from here. The assumption that a new flagship shows up in your tool’s model picker within a week no longer holds.

Buried further down the same page is the disclosure that deserves the most attention and has received the least: Astra’s written reasoning is harder to monitor than Sol’s. OpenAI attributes this to Astra solving problems in fewer written steps, describes the decline as serious, and says monitorability remains a research priority.

If any part of your governance story is “we can read what the agent was thinking,” read that sentence twice. The trend line on chain-of-thought visibility is pointing the wrong way, and this is the first flagship where the vendor has said so plainly.

In fairness, the safety news isn’t all bad: in OpenAI’s honeypot tests Astra stayed inside its authorised scope in every case, where Sol strayed 48.2% of the time. Better behaved, harder to audit. Those aren’t contradictory, but they’re an uncomfortable pair.

How this reaches your stack

If you use Cursor: it doesn’t. OpenAI is pulling its models out of Cursor on November 12 and has said it won’t supply new ones in the meantime, so Astra never appears in Cursor’s picker at all. We covered that in OpenAI is pulling its models out of Cursor.

If you use Codex: this is where Astra lands first and where it makes most sense. OpenAI updated the Codex harness alongside the launch and claims 1.9x faster task completion on Mind2Web. There’s also an experimental feature worth knowing about: in long sessions Astra can keep notes across context windows rather than compressing everything into one summary each time the context fills, and earlier windows stay searchable. You switch it on in your Codex config. See our Codex review.

If you build with Lovable, Bolt, Replit, or Base44: you will never see “Astra” as a setting, and given the pricing and the gating, there’s little reason for these tools to route to it. They optimise for cost per generated app, and Astra is expensive at a capability that app builders don’t use.

If you use Claude Code: nothing changes. Fable 5.1 remains ahead on the independent index at the same price.

What to actually do

If you run browser or desktop automation, get in the queue. This is the one clear case. A model that completes a multi-step job inside real software at half the time per task changes the build-versus-drive maths for anyone currently hand-writing connectors to internal tools. Pilot it on a workload you can afford to have interrupted — the safety checks can slow, pause, or stop legitimate work, and in the API a flagged task simply stops.

If you’re building an app, do nothing. Your tool’s model choice is not your bottleneck and hasn’t been for a year. Astra is level with what you already have access to, at the same price, with less availability.

Stop treating flagship launches as events. Four frontier labs shipped inside 72 hours in early September. The lead rotates every few weeks, the benchmarks are vendor-run, and the gaps are inside the noise. Pick tools that absorb whichever model is best this month, and keep your review process tight regardless of what’s underneath. That advice is in picking AI tools that last and this launch is the strongest evidence for it yet.

The genuinely new things this month were a model gated for security reasons and a vendor admitting its model got harder to monitor. Neither of those is a benchmark, and both will matter longer than the benchmarks will.

Source: OpenAI’s GPT-6 Astra announcement, with benchmark and rollout detail via DataNorth.

Related guides

founder · 8 min read

35 Security Holes in One Month: Why Vibe-Coded Apps Are Getting Riskier in 2026

35 new CVEs in March 2026 were traced to AI-generated code. Here's what happened and what founders need to do about it.

securityvibe coding
Mar 2026

founder · 6 min read

Apple Just Made AI Free to Put in Your App. What WWDC 2026 Means for Founders.

Apple's WWDC 2026 gave founders free on-device AI, a free cloud tier, and an agentic Xcode. Here's what actually matters if you're shipping a mobile app.

applewwdc
Jun 2026

founder · 6 min read

AWS Just Entered the Vibe-Coding Race. Here's What Founders Should Take From It.

AWS and Superblocks signed a multiyear deal to run governed vibe coding inside private clouds. The real signal isn't the product — it's who's moving.

awssuperblocks
Aug 2026

Recommended next step

Was this helpful?