OpenAI released GPT-6 Astra on 3 September 2026, to limited customers first. Co-founder Greg Brockman said it would be "reasonable to see Astra as a version of artificial general intelligence" and opened with "Welcome to the AGI era!"
The launch table carries a number that has done most of the talking since: 98.6% on ARC-AGI-3, against 7.8% for GPT-5.6 Sol. That is a twelve-fold jump inside one vendor's own generation, on the benchmark most associated with the phrase "general intelligence."
It is also the number that needs reading most carefully, and OpenAI has already published the reason why.
Everything below is OpenAI's, about OpenAI's model. None of it was reproduced here, and the columns for Claude and Gemini are a rival's measurements of a rival's models — the least trustworthy kind of figure on any launch page.
The launch table
| evaluation | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| ARC-AGI-3 | 98.6% | 7.8% | — | — | 30.2% (high) | — |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 87.8% | 73.2% | — |
| Agents' Last Exam | 59.3% | 52.7% | — | 48.7% (xhigh) | 52.7% | — |
| AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% | — |
| BenchCAD | 95.9% | 83.3% | 84.3% | 67.5% | 82.1% | — |
| DeepSWE v1.1 | 74.1% | 70.8% | 67.4% | 69.9% | 68.8% | 73.7% |
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 24.7% | 29.0% | — |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.2% | 95.3% |
| GeneBench Pro | 39.0% | 28.7% | — | — | — | — |
| MedChemBench (internal) | 49.7% | 47.4% | — | — | — | — |
| HealthBench Professional | 63.4% | 60.5% | 56.6% | 60.9% | 57.5% | 52.1% |
| ExploitBench | 100.0% | 78.5% | — | — | 70% | — |
| SRE-Bench (four attempts) | 99.2% | 68.7% | — | — | — | — |
| Auto-review circumvention (lower is better) | 0% | 0.29% | — | — | — | — |
Note how many cells are empty. On six of the fourteen rows there is no competitor number at all, so "Astra leads" on those rows is unopposed by construction rather than by measurement.
What the 98.6% actually measures
In July, OpenAI published a post with a title that reads differently today: "How enabling two settings tripled our scores on the ARC-AGI-3 benchmark."
The two settings were retained reasoning between turns and compaction — both features of its Responses API, neither a change to the model. Reimplementing the ARC-AGI-3 harness that way took GPT-5.6 Sol from 13.3% to 38.3% on the public task set, and cut output tokens by a factor of six. The explanation OpenAI gave is that the model stopped re-deriving each game's rules from scratch and held on to strategies it had already worked out.
Its own conclusion, in its own words: benchmark results say as much about harness design as they do about the underlying model.
So there are now three published OpenAI figures for GPT-5.6 Sol on ARC-AGI-3 — 7.8%, 13.3% and 38.3% — and the launch table pairs Astra's 98.6% against the lowest of them. They may not be the same measurement; the July post specifies the public task set and the launch table does not say. That is precisely the problem. A twelve-fold gap that includes an unstated harness difference is not a clean generational comparison, and the company that published the caveat is the one now quoting the number.
There is now independent corroboration, and it is blunt. VentureBeat's coverage notes that NVIDIA's AVO architecture reached 100% on ARC-AGI-3 using Claude Opus 5 as its base model — the same Claude Opus 5 that appears at 30.2% in OpenAI's own launch table. What closed a 70-point gap was not a better model but persistent memory, tools, feedback and recovery wrapped around one.
If a rival's model can be driven from 30% to 100% on this benchmark by changing the scaffolding, then ARC-AGI-3 is substantially a measurement of scaffolding. That is the same conclusion OpenAI reached in July about its own results, arrived at independently by a third party using a competitor's weights.
None of which means Astra is not a large step. It means 98.6% is a score for Astra plus an agent system, and the honest version of the claim names the harness.
The part that is not a benchmark
The more consequential news is not in the table.
On 2 September, Astra crossed the Critical cybersecurity capability threshold in OpenAI's Preparedness Framework — the first model to do so. OpenAI says it "can find previously unknown security flaws and develop ways to exploit them" across protected systems without human guidance. ExploitBench: 100.0%.
OpenAI says Astra discovered two previously unknown vulnerabilities during evaluation. The response was procedural rather than rhetorical. Increased cybersecurity protocols over the preceding weeks, increased monitoring so OpenAI can "rapidly detect and contain potentially misaligned actions", and — the detail that matters most for anyone hoping to use it — Astra's advanced cyber capabilities shipped first through Daybreak Blue, prioritising critical infrastructure defenders, with general access carrying "stronger restrictions and monitoring" and wider enterprise and consumer availability following.
Which should sound familiar if you read what Anthropic shipped days earlier. Fable 5.1 and Mythos 5.1 arrived described as the strongest cyber capabilities of any released model, with roughly 60% fewer safeguard interventions per session, vulnerability identification newly permitted while exploit generation stays refused — and Mythos reserved for a Cyber Verification Program for defensive security work.
Two labs, one week. One tightened after crossing a line it had drawn; the other loosened before reaching one. Both ended at the same place: frontier cyber capability behind a defenders-only door. That is now the industry's default answer, and it happened without anyone announcing it as a policy.
Where OpenAI's Claude numbers meet Anthropic's
Two benchmarks appear in both this table and Anthropic's own launch table from two days earlier, which makes them checkable.
Terminal-Bench Science 0.1 agrees exactly, on all four shared cells: Fable 5.1 52.6, Fable 5 24.7, Opus 5 29.0, GPT-5.6 Sol 22.4. Four independent runs matching to a tenth of a point does not happen. One lab is citing the other's published figures.
AutomationBench does not:
| Anthropic published | OpenAI published | |
|---|---|---|
| Claude Fable 5.1 | 31.4% | 31.4% |
| Claude Fable 5 | 17.1% | 17.4% |
| Claude Opus 5 | 26.9% | 26.9% |
| GPT-5.6 Sol | 19.6% | 18.1% |
The interesting part is the direction. Each lab scores its rival's model higher than the rival does. Anthropic put GPT-5.6 Sol 1.5 points above OpenAI's own figure; OpenAI put Claude Fable 5 0.3 points above Anthropic's. If either were massaging the comparison, this is the wrong way round.
The dull explanation is almost certainly the right one — different harnesses, different runs, ordinary variance. But it is worth knowing that two vendors publishing the same benchmark in the same week produced different numbers for the same models, and that the gaps are the size of the gaps launch posts routinely treat as decisive.
The benchmark that is missing
One absence is worth as much as any number present. OpenAI did not publish GDPval results for Astra — its own 2025 benchmark, built to measure performance on economically valuable real-world tasks across 44 knowledge-work occupations.
That is a strange thing to leave out of an AGI announcement, since GDPval is the OpenAI benchmark most directly aimed at the question "can this do the work a professional does." It is also the gap that shows up most clearly against the competition: Anthropic did publish a GDPval-AA v2 column two days earlier — 1853 for Fable 5.1 against 1824 for Opus 5 and 1711 for GPT-5.6 Sol.
So the benchmark designed to ask whether a model can do real work is the one the AGI-era launch skips, and the rival's number for OpenAI's previous model is the only one in public.
Price, and what it costs to be first
| model ID | gpt-6-astra |
| standard | $10 / MTok input, $50 / MTok output |
| fast mode | $20 / $100 per MTok |
Which is, to the dollar, what Anthropic charges for Claude Fable 5.1 — $10 and $50. Two labs, two weeks apart, arriving at identical list prices for their frontier tier.
OpenAI's efficiency claim is the more interesting number: roughly 57% lower estimated API cost per task than GPT-5.6 Sol's highest-scoring configuration on DeepSWE v1.1. And on OSWorld 2.0's offline subset it reports 72.6% at about 40 minutes per task, against Sol's 65.7% at about 75 minutes — a rare launch figure that quotes wall clock alongside the score, which is the honest way to report agentic work.
What this does not tell you
Nothing here was run by us. Every figure is OpenAI's, including all the competitor columns. We have no access to Astra.
Six of fourteen rows have no competitor at all. GeneBench Pro, MedChemBench, SRE-Bench, auto-review circumvention and the two ARC-AGI-3 Claude cells are blank or partial. A lead over an empty column is not a lead.
Two of the evaluations are internal — MedChemBench and auto-review circumvention are labelled as OpenAI's own, so there is no external definition to check them against.
The disagreement is live, not settled. Commentary around the launch questions whether ARC-AGI-3 results reflect the foundation model or the whole agent system, and whether a result like NVIDIA's represents generalisation or overfitting to one benchmark. The benchmark's context restrictions have also been called an unrealistic model of how production agents actually run.
And "AGI" is a claim by an interested party. Brockman's framing is a position, not a measurement, and the benchmark most cited in support of it is the one whose harness dependence OpenAI documented itself.
Related reading
- Claude Fable 5.1 vs Fable 5.0: What Changed, What Breaks — the table this one is checked against
- Claude Fable 5.1 and Mythos 5.1: the Science Results — the other lab's cyber week, and its defenders-only programme
- Claude Mythos Preview: Why Anthropic Locked Its Best Security Model Behind a Wall — the first time a lab gated a model on security grounds
- Best Model for a Dual DGX Spark — what happens when you measure the workload instead of the benchmark