AI News 9 min read

GPT-6 Astra Benchmarks: What the 98.6% on ARC-AGI-3 Actually Measures

ai.rs Sep 3, 2026
GPT-6 Astra Benchmarks: What the 98.6% on ARC-AGI-3 Actually Measures illustration

OpenAI released GPT-6 Astra on 3 September 2026, to limited customers first. Co-founder Greg Brockman said it would be "reasonable to see Astra as a version of artificial general intelligence" and opened with "Welcome to the AGI era!"

The launch table carries a number that has done most of the talking since: 98.6% on ARC-AGI-3, against 7.8% for GPT-5.6 Sol. That is a twelve-fold jump inside one vendor's own generation, on the benchmark most associated with the phrase "general intelligence."

It is also the number that needs reading most carefully, and OpenAI has already published the reason why.

Everything below is OpenAI's, about OpenAI's model. None of it was reproduced here, and the columns for Claude and Gemini are a rival's measurements of a rival's models — the least trustworthy kind of figure on any launch page.

The launch table

evaluation GPT-6 Astra GPT-5.6 Sol Claude Fable 5.1 Claude Fable 5 Claude Opus 5 Gemini 3.8 Flash
ARC-AGI-3 98.6% 7.8% 30.2% (high)
FrontierMath Tier 4 (v2) 97.6% 83.0% 87.8% 87.8% 73.2%
Agents' Last Exam 59.3% 52.7% 48.7% (xhigh) 52.7%
AutomationBench 41.4% 18.1% 31.4% 17.4% 26.9%
BenchCAD 95.9% 83.3% 84.3% 67.5% 82.1%
DeepSWE v1.1 74.1% 70.8% 67.4% 69.9% 68.8% 73.7%
Terminal-Bench Science 0.1 64.6% 22.4% 52.6% 24.7% 29.0%
GPQA Diamond 96.0% 94.6% 93.7% 92.6% 93.2% 95.3%
GeneBench Pro 39.0% 28.7%
MedChemBench (internal) 49.7% 47.4%
HealthBench Professional 63.4% 60.5% 56.6% 60.9% 57.5% 52.1%
ExploitBench 100.0% 78.5% 70%
SRE-Bench (four attempts) 99.2% 68.7%
Auto-review circumvention (lower is better) 0% 0.29%

Note how many cells are empty. On six of the fourteen rows there is no competitor number at all, so "Astra leads" on those rows is unopposed by construction rather than by measurement.

What the 98.6% actually measures

GPT-5.6 Sol scored 7.8, 13.3 or 38.3 percent on ARC-AGI-3 depending on the harness, before Astra's 98.6

In July, OpenAI published a post with a title that reads differently today: "How enabling two settings tripled our scores on the ARC-AGI-3 benchmark."

The two settings were retained reasoning between turns and compaction — both features of its Responses API, neither a change to the model. Reimplementing the ARC-AGI-3 harness that way took GPT-5.6 Sol from 13.3% to 38.3% on the public task set, and cut output tokens by a factor of six. The explanation OpenAI gave is that the model stopped re-deriving each game's rules from scratch and held on to strategies it had already worked out.

Its own conclusion, in its own words: benchmark results say as much about harness design as they do about the underlying model.

So there are now three published OpenAI figures for GPT-5.6 Sol on ARC-AGI-3 — 7.8%, 13.3% and 38.3% — and the launch table pairs Astra's 98.6% against the lowest of them. They may not be the same measurement; the July post specifies the public task set and the launch table does not say. That is precisely the problem. A twelve-fold gap that includes an unstated harness difference is not a clean generational comparison, and the company that published the caveat is the one now quoting the number.

There is now independent corroboration, and it is blunt. VentureBeat's coverage notes that NVIDIA's AVO architecture reached 100% on ARC-AGI-3 using Claude Opus 5 as its base model — the same Claude Opus 5 that appears at 30.2% in OpenAI's own launch table. What closed a 70-point gap was not a better model but persistent memory, tools, feedback and recovery wrapped around one.

If a rival's model can be driven from 30% to 100% on this benchmark by changing the scaffolding, then ARC-AGI-3 is substantially a measurement of scaffolding. That is the same conclusion OpenAI reached in July about its own results, arrived at independently by a third party using a competitor's weights.

None of which means Astra is not a large step. It means 98.6% is a score for Astra plus an agent system, and the honest version of the claim names the harness.

The part that is not a benchmark

The more consequential news is not in the table.

OpenAI crossed a Critical cyber threshold and gated Astra behind a defenders programme; Anthropic loosened safeguards and gated Mythos behind one

On 2 September, Astra crossed the Critical cybersecurity capability threshold in OpenAI's Preparedness Framework — the first model to do so. OpenAI says it "can find previously unknown security flaws and develop ways to exploit them" across protected systems without human guidance. ExploitBench: 100.0%.

OpenAI says Astra discovered two previously unknown vulnerabilities during evaluation. The response was procedural rather than rhetorical. Increased cybersecurity protocols over the preceding weeks, increased monitoring so OpenAI can "rapidly detect and contain potentially misaligned actions", and — the detail that matters most for anyone hoping to use it — Astra's advanced cyber capabilities shipped first through Daybreak Blue, prioritising critical infrastructure defenders, with general access carrying "stronger restrictions and monitoring" and wider enterprise and consumer availability following.

Which should sound familiar if you read what Anthropic shipped days earlier. Fable 5.1 and Mythos 5.1 arrived described as the strongest cyber capabilities of any released model, with roughly 60% fewer safeguard interventions per session, vulnerability identification newly permitted while exploit generation stays refused — and Mythos reserved for a Cyber Verification Program for defensive security work.

Two labs, one week. One tightened after crossing a line it had drawn; the other loosened before reaching one. Both ended at the same place: frontier cyber capability behind a defenders-only door. That is now the industry's default answer, and it happened without anyone announcing it as a policy.

Where OpenAI's Claude numbers meet Anthropic's

Two benchmarks appear in both this table and Anthropic's own launch table from two days earlier, which makes them checkable.

Terminal-Bench Science 0.1 agrees exactly, on all four shared cells: Fable 5.1 52.6, Fable 5 24.7, Opus 5 29.0, GPT-5.6 Sol 22.4. Four independent runs matching to a tenth of a point does not happen. One lab is citing the other's published figures.

AutomationBench does not:

Anthropic published OpenAI published
Claude Fable 5.1 31.4% 31.4%
Claude Fable 5 17.1% 17.4%
Claude Opus 5 26.9% 26.9%
GPT-5.6 Sol 19.6% 18.1%

The interesting part is the direction. Each lab scores its rival's model higher than the rival does. Anthropic put GPT-5.6 Sol 1.5 points above OpenAI's own figure; OpenAI put Claude Fable 5 0.3 points above Anthropic's. If either were massaging the comparison, this is the wrong way round.

The dull explanation is almost certainly the right one — different harnesses, different runs, ordinary variance. But it is worth knowing that two vendors publishing the same benchmark in the same week produced different numbers for the same models, and that the gaps are the size of the gaps launch posts routinely treat as decisive.

The benchmark that is missing

One absence is worth as much as any number present. OpenAI did not publish GDPval results for Astra — its own 2025 benchmark, built to measure performance on economically valuable real-world tasks across 44 knowledge-work occupations.

That is a strange thing to leave out of an AGI announcement, since GDPval is the OpenAI benchmark most directly aimed at the question "can this do the work a professional does." It is also the gap that shows up most clearly against the competition: Anthropic did publish a GDPval-AA v2 column two days earlier — 1853 for Fable 5.1 against 1824 for Opus 5 and 1711 for GPT-5.6 Sol.

So the benchmark designed to ask whether a model can do real work is the one the AGI-era launch skips, and the rival's number for OpenAI's previous model is the only one in public.

Price, and what it costs to be first

model ID gpt-6-astra
standard $10 / MTok input, $50 / MTok output
fast mode $20 / $100 per MTok

Which is, to the dollar, what Anthropic charges for Claude Fable 5.1 — $10 and $50. Two labs, two weeks apart, arriving at identical list prices for their frontier tier.

OpenAI's efficiency claim is the more interesting number: roughly 57% lower estimated API cost per task than GPT-5.6 Sol's highest-scoring configuration on DeepSWE v1.1. And on OSWorld 2.0's offline subset it reports 72.6% at about 40 minutes per task, against Sol's 65.7% at about 75 minutes — a rare launch figure that quotes wall clock alongside the score, which is the honest way to report agentic work.

What this does not tell you

Nothing here was run by us. Every figure is OpenAI's, including all the competitor columns. We have no access to Astra.

Six of fourteen rows have no competitor at all. GeneBench Pro, MedChemBench, SRE-Bench, auto-review circumvention and the two ARC-AGI-3 Claude cells are blank or partial. A lead over an empty column is not a lead.

Two of the evaluations are internal — MedChemBench and auto-review circumvention are labelled as OpenAI's own, so there is no external definition to check them against.

The disagreement is live, not settled. Commentary around the launch questions whether ARC-AGI-3 results reflect the foundation model or the whole agent system, and whether a result like NVIDIA's represents generalisation or overfitting to one benchmark. The benchmark's context restrictions have also been called an unrealistic model of how production agents actually run.

And "AGI" is a claim by an interested party. Brockman's framing is a position, not a measurement, and the benchmark most cited in support of it is the one whose harness dependence OpenAI documented itself.

Frequently Asked Questions

What is GPT-6 Astra? +

OpenAI's model released on 3 September 2026, initially to limited customers. Co-founder Greg Brockman said it would be "reasonable to see Astra as a version of artificial general intelligence" and described it as a jump in capability. OpenAI's launch materials call it the world's best computer use model, designed to operate browsers, spreadsheets and desktop applications and carry out multistep workflows. It shipped first to the Daybreak programme for cybersecurity defenders, with wider enterprise and consumer access planned.

Did GPT-6 Astra really score 98.6% on ARC-AGI-3? +

That is the figure OpenAI published, but the harness matters and OpenAI has documented that itself. In July it published a post titled "How enabling two settings tripled our scores on the ARC-AGI-3 benchmark", showing that retained reasoning between turns plus compaction — both Responses API features, neither a model change — took GPT-5.6 Sol from 13.3% to 38.3% on the public task set while cutting output tokens sixfold. Its stated conclusion was that benchmark results say as much about harness design as about the underlying model. The 98.6% is a score for Astra plus an agent system, and the launch table does not name the harness.

Why does GPT-5.6 Sol have three different ARC-AGI-3 scores? +

Because they were measured differently. OpenAI's launch table cites 7.8%; its own July run with the official harness gave 13.3%; and with retained reasoning and compaction enabled the same model scored 38.3% on the public task set. The July post specifies the public set and the launch table does not, so the figures may not be comparable — which is the problem. Astra's 98.6% is presented against the lowest of the three.

Is GPT-6 Astra dangerous? +

OpenAI says Astra crossed the Critical cybersecurity capability threshold in its Preparedness Framework on 2 September, the first model to do so, and that it can find previously unknown security flaws and develop ways to exploit them without human guidance. It scores 100% on ExploitBench. OpenAI's response was to increase cybersecurity protocols and monitoring so it can rapidly detect and contain potentially misaligned actions, and to ship Astra first to a programme for cybersecurity defenders rather than to general availability.

Does the ARC-AGI-3 score measure the model or the harness? +

Substantially the harness, on the available evidence. OpenAI published in July that two Responses API settings — retained reasoning between turns and compaction — took GPT-5.6 Sol from 13.3% to 38.3% on the public task set with no model change, concluding that benchmark results say as much about harness design as about the model. Independently, NVIDIA's AVO architecture is reported to have reached 100% on ARC-AGI-3 using Claude Opus 5 as its base — the same model OpenAI's launch table lists at 30.2% — by adding persistent memory, tools, feedback and recovery. If scaffolding can move a model from 30% to 100%, the benchmark is largely measuring scaffolding.

How does GPT-6 Astra compare to Claude Fable 5.1? +

On OpenAI's own table Astra leads on every shared row, but those Claude figures are a rival's measurements. Two benchmarks can be checked against Anthropic's launch table from two days earlier: Terminal-Bench Science 0.1 agrees exactly on all four shared cells, which suggests one lab cited the other rather than re-running it, while AutomationBench differs — Anthropic reports Fable 5 at 17.1% against OpenAI's 17.4%, and GPT-5.6 Sol at 19.6% against OpenAI's own 18.1%. Notably each lab scores the rival's model higher than the rival does.

Can I use GPT-6 Astra? +

Not immediately for most people. It launched on 3 September to limited customers, starting with participants in OpenAI's Daybreak programme for cybersecurity defenders, with wider access for enterprise and consumer accounts said to be coming in the following days.

How much does GPT-6 Astra cost? +

The model ID is gpt-6-astra, priced at $10 per million input tokens and $50 per million output tokens, with a fast mode at $20 and $100. That is identical to Anthropic's list price for Claude Fable 5.1. OpenAI also claims roughly 57% lower estimated API cost per task than GPT-5.6 Sol's highest-scoring configuration on DeepSWE v1.1, which is a claim about efficiency at a fixed task rather than about the per-token rate.

What does this mean for your business?

New models drop every month. The real question is whether the underlying capability fits your business. Find out in 2 minutes.

Take the AI Readiness Check
Share: Post Share

Read next