5 SEP 2026 — OpenAI released GPT-6 Astra on 3 September, describing it as its most capable model for complex work and its best model for software engineering to date. President Greg Brockman called the launch the start of the AGI era. On the one general index that OpenAI neither ran nor selected, Astra scores exactly what GPT-5.6 Sol scores, at two and a half times the price per token.
What shipped, and who can actually use it
Astra began rolling out on 3 September to a limited set of organisations, with OpenAI saying it would reach all ChatGPT Plus, Pro, Business and Enterprise users over the following days, along with the OpenAI API and AWS. Enterprise workspaces have it switched off by default, so an administrator has to enable it before anyone inside the company sees it.
API pricing is 10 dollars per million input tokens and 50 dollars per million output, up from 4 and 20 for GPT-5.6 Sol. Cached input is a dollar, batch runs at half price, and a fast mode charges double. The context window is a million tokens.
The launch itself was untidy. Several outlets published from OpenAI's own launch material while the company's Astra page was returning errors to ordinary visitors, and archived mirrors circulated for close to an hour before the page came up properly.
This lands four weeks after OpenAI paused the model for two weeks over cyber risk, and its cybersecurity results, including a perfect ExploitBench run against a target built for testing, are covered separately and not repeated here.
Level with the model it replaces
Artificial Analysis, which publishes its methodology and runs its evaluations without the model providers involved, is the closest thing to a neutral scoreboard for this launch. On its Intelligence Index, Astra scores 61 — the same as GPT-5.6 Sol.
Claude Fable 5.1 leads that index at 66 and Claude Opus 5 sits at 63, with Meta's Muse Spark 1.3 also ranking above Astra. Astra does better on the same house's Coding Agent Index, at 67 in the Codex harness, roughly level with Opus 5 and Claude Fable 5, while Fable 5.1 running in Claude Code leads at 70.
Put that score next to the price. Astra uses about a tenth fewer tokens than Sol at maximum effort, which does not come close to absorbing a 2.5x price rise, so the same evaluation runs about 75 per cent more expensive per task. Coding is the exception: it uses roughly a third of Sol's tokens and lands at about the same cost for two more points.
None of this makes Astra a bad model. It does leave a gap between the launch framing and the neutral measurement, and a buyer should notice that gap before signing anything.
The software engineering claim is a tie
OpenAI's own published figures argue against its best-model claim. On DeepSWE v1.1, Astra scores 74.1, Gemini 3.8 Flash 73.8 and Claude Opus 5 73.7 — a spread of four tenths of a point across three labs. Meta separately reported 75.4 for Muse Spark 1.3, which is higher than any of them.
On FrontierCode 1.1 the ordering reverses outright. Astra at 53.3 trails Claude Fable 5 at 53.5 and Opus 5 at 53.4. Those are differences well inside the range where re-running the same test on a different day moves the ranking.
Terminal-Bench 4.0 is the one coding benchmark with a real gap, and Astra leads it at 57.7 against Fable 5.1 at 55.8. One clear win and two statistical ties is a good result, and it is not the best model for software engineering to date.
The 99.9 per cent is a harness
The number doing the most work in the coverage is ARC-AGI-3, where Astra reached 99.9 per cent against Sol's 7.8. That figure was produced with a stateful, expensive harness that OpenAI built. On ARC's standard harness the same model scores 66 per cent.
The two results are not contradictory; they measure different things. The 66 per cent is what a developer calling the API gets, and it is the figure that has largely not travelled.
The same caution applies to FrontierMath Tier 4, where Astra's 97.6 per cent leads Fable 5.1's 87.8. That is the widest academic gap in the set. OpenAI funded FrontierMath's development and held exclusive access to part of it; that was disclosed, and it is worth remembering when reading the margin.
The announcement prose omits one academic benchmark. On Humanity's Last Exam with tools, Astra scores 57.2 against Fable 5.1 at 65.0, Fable 5 at 63.8 and Opus 5 at 63.6. It is the only academic test in the published set where Astra comes last, and the silence around it is the clearest signal in the release about which numbers were chosen.
Where Astra separates is time, not score
Behind the contested benchmarks, something solid remains. On OSWorld 2.0, which tests real work across desktop applications, Astra scores 72.6 against Sol's 65.7 — though Opus 5 is at 70.2, so the lead over a rival is 2.4 points rather than the seven the announcement implies. On ScreenSpot-Pro it scores 92.7 against Fable 5's 87.3, a lead large enough to matter.
The figure that deserved the headline is the clock. Average time per OSWorld task fell from about 75 minutes to about 40, and on Mind2Web the model completes tasks 1.9 times faster than its predecessor. On Agents' Last Exam it beats Opus 5 while using roughly 65 per cent fewer output tokens.
That is a different and more useful claim than a higher score. A model that finishes an hour of desktop work in forty minutes changes what can be delegated to it, because the constraint on agentic work has never been whether the model eventually gets there. It is whether anyone will wait.
What would settle it
Three disclosures, none of which accompanied this launch. The harness used for every headline figure, stated next to the figure. Whether a benchmark's development was funded by the lab reporting the score. And a result on a suite the lab did not select, published at the same time rather than by a third party days later.
Until then, Astra looks like a strong computer-use model at a much higher price, sold with an AGI framing that its own coding tables do not support. We wrote yesterday that the week's release count was inflated. The benchmark count is inflated the same way, and for the same reason: whoever picks the test picks the winner.