Full table ARC-AGI-3 OSWorld Independent Missing Head-to-head Checklist FAQ
GPT-6 Astra · Benchmark Teardown

GPT-6 Astra Benchmarks Explained: Every Score, and Why the Records Have Asterisks

GPT-6 Astra sets records on what it was built for and is flat on general intelligence. But the headline ARC-AGI-3 number swings 37 points depending on the harness, and rival models were scored under different setups. The asterisks matter more than the scores.

GPT-6 Astra benchmark teardown — comparing scores across harnesses
The scores are fine. The harnesses are the story.

Quick share

For anyone about to quote the 99.9% figure.

Get a summary from AI

Short on time? Open this article in an answer engine and have it summarised for you.

The one-paragraph verdict

GPT-6 Astra sets records on the things it was built for — computer use, terminals, CAD, cyber — and is statistically flat on general intelligence. The headline ARC-AGI-3 number is 99.9% under OpenAI's own Provider Adapter harness and 62.7% on the standard one, and rival models were scored under different setups. The OSWorld comparison against Claude Fable 5.1 uses a different version of the test. GDPval is missing entirely. None of it has been independently replicated. It is a genuinely strong specialised model with an unusually noisy scoreboard.

This is the companion teardown to the full GPT-6 Astra guide. The pillar covers what the model is; this page covers what the numbers mean, which ones are comparable, and which ones quietly aren't.

The reason it needs its own page: most launch coverage reprinted the scores in a table and moved on. The scores are fine. The harnesses are the story.

1. The full benchmark table

Everything OpenAI published, grouped by family, with comparisons where they exist and are meaningful.

Computer use — the category Astra was built to win

BenchmarkAstraGPT-5.6 SolNotes
OSWorld 2.072.6%65.7%~40 min/task vs ~75 min — see §3
Agents' Last Exam59.3%Claude Opus 5: 55.5%
ScreenSpot-Pro92.7%UI element grounding

The time-per-task figure matters more than the accuracy figure here. A 6.9-point accuracy gain is respectable; doing it in 47% less wall-clock time is what changes whether an agent workflow is economically viable.

Coding and engineering

BenchmarkAstraComparison
Terminal-Bench 4.057.9%Sol: 37.3%
DeepSWE v1.174.1%Sol: 70.8%; Meta Muse 1.3 reportedly 75.4%
FrontierCode Extended64.5%

Terminal-Bench is the standout — a 20-point jump is not incremental. Note also that on DeepSWE, the benchmark OpenAI leans on hardest for its cost-per-task argument, Astra is behind a reported Meta Muse score. That comparison did not feature prominently in the launch materials.

Science and reasoning

BenchmarkAstraComparison
FrontierMath Tier 4 v297.6%
GPQA Diamond96.0%
ARC-AGI-399.9% / 62.7%Depends entirely on harness — see §2
Terminal-Bench Science 0.164.6%Claude Fable 5.1: 52.6%
BenchCAD Vision2Code95.9%Fable 5.1: 84.3%
Humanity's Last Exam57.2%

GPQA Diamond at 96% is effectively saturated — when a benchmark approaches its ceiling it stops discriminating between models, and should be retired from comparison rather than quoted as a win.

Cybersecurity

BenchmarkAstraGPT-5.6 Sol
ExploitBench100%78.5%
ExploitGym42.4%30.3%
SRE-Bench (single attempt)88.0%

A 100% score should always prompt the same question: is the benchmark exhausted, or is the model that good? ExploitGym at 42.4% — harder, more open-ended — is the more informative number, and it is the one that justifies the Critical classification alongside the two zero-days found during evaluation.

Alignment and safety

MetricAstraPrior
Safety benchmark error rate2.4%
Circumvention attempts0%Sol: 48.2–56%
Hallucination benchmark error rate4.2%
Cyber jailbreak rejection91.5%59%

2. The ARC-AGI-3 problem

This is the number that made the headlines and the one that survives scrutiny least well.

Provider Adapter harness99.9%
Standard harness62.7%
Gap37.2 points, same model

OpenAI ran Astra through its own Responses API harness with two settings changed "to better reflect how the model performs in real-world use." The Provider Adapter, as Simon Willison flagged, "preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work."

Three things follow:

  1. It is a real engineering result. Persisting reasoning state across requests genuinely helps, and it reflects how you would actually deploy the model. Dismissing the number entirely is as wrong as quoting it uncritically.
  2. It is not a like-for-like comparison. The rival models in the same chart ran under different setups. A 37-point swing from harness configuration is larger than the gap between most frontier models.
  3. Reported figures vary. Some outlets cite 98.6% rather than 99.9%, likely different runs or settings. The variance itself is informative.

ARC Prize was unusually direct about the interpretation: "we are not claiming that it is AGI," noting the benchmark has "deterministic, closed-ended mechanics" that don't represent real-world complexity. Greg Kamradt's more interesting finding got far less coverage: on ARC-AGI-3, Astra surpassed the human action-efficiency baseline on 96% of levels — a claim about solving unfamiliar environments in fewer moves than people, which is a different and arguably more impressive thing than raw accuracy.

The rule this establishes

A benchmark score without its harness is not a number, it's a marketing asset. From now on, any frontier score you cite should carry the harness, the settings and whether competitors ran the same configuration. If a vendor won't publish that, treat the score as unverified.

3. The OSWorld version mismatch

The second comparability problem, and a quieter one.

Astra scores 72.6% on OSWorld 2.0. Anthropic reported 77.9% for Claude Fable 5.1 on OSWorld — a higher number, on a different version of the test.

Read carelessly, that says Fable beats Astra at computer use. Read carefully, it says nothing at all: benchmark versions change task sets, scoring and difficulty. You cannot subtract one from the other.

What you can say with confidence is the internally consistent comparison: Astra 72.6% vs GPT-5.6 Sol 65.7% on the same version, at roughly half the time per task. That is a real generational gain within one vendor's own lineage — and it is the claim OpenAI is entitled to make.

4. Artificial Analysis: the independent read

The most useful third-party evidence available, because they run their own harness across every model — which is exactly the property the vendor numbers lack.

MeasureResultWhat it means
Intelligence Index61Level with GPT-5.6 Sol; behind Claude Fable 5.1
Coding67Matching Claude Opus 5 and Fable 5
Token efficiency~1/3 of Sol's tokensIn the Codex harness, at roughly equal cost, for better results
Hallucination rate92% → 51% at max effortAccuracy up 4 points alongside it
Economic / support tasksRegressionsMeasurably worse than the predecessor

Two of those deserve more attention than they got.

The hallucination collapse is the most underrated number in the launch. A drop from 92% to 51% on their measure, with accuracy improving simultaneously, is a larger practical improvement for most real workloads than any computer-use record. Fewer confident wrong answers changes how much review a human has to do, which is the actual cost driver in production.

The regressions matter more than the records for most buyers. If your workload is customer support, this model is measurably worse than its predecessor at 2.5× the price. That is the single most actionable finding on this page.

Their conclusion — Astra is a specialised model, not a universally superior one — is the honest summary of the whole scoreboard.

5. What's missing

GPT-6 AstraClaude Fable 5.1GPT-5.6 SolMeta Muse
Price per 1M in/out$10 / $50$10 / $50$4 / $20$1.25 / $4.25 (Spark)
AA Intelligence Index61Ahead of Astra~61
Computer use72.6% OSWorld 2.077.9% on a different version65.7%
DeepSWE v1.174.1%70.8%75.4% reported (1.3)
Science terminal64.6%52.6%
Vision-to-code (BenchCAD)95.9%84.3%

The summary that survives all the caveats: Astra and Fable 5.1 cost exactly the same and win different things. Astra takes CAD, science terminals and cyber. Fable takes general intelligence. Sol remains 60% cheaper and is not obviously worse for chat and general reasoning. Muse is an order of magnitude cheaper and competitive on the one coding benchmark where a comparison exists.

That is not a leaderboard. It is a routing table — and routing is what you should be building, rather than picking a winner.

7. How to read any frontier benchmark from now on

The six questions

  • What harness produced this number? If it's the vendor's own, was it the same one used for the competitors in the chart?
  • What version of the benchmark? OSWorld 2.0 and OSWorld are not the same test.
  • Is the benchmark saturated? Anything above ~95% has stopped discriminating.
  • Which results are absent? A missing benchmark from a vendor that owns it is a data point.
  • Has anyone independent reproduced it? Vendor numbers are a hypothesis until then.
  • Was failure included? Clean runs tell you nothing about behaviour under tool failures, interruptions or ambiguous permissions.

Apply those six to any launch and most scoreboards get considerably quieter. That is the correct outcome.

8. Frequently asked questions

Did GPT-6 Astra really score 99.9% on ARC-AGI-3?

Under OpenAI's own Provider Adapter harness, yes. On the standard harness the same model scores 62.7%, and competing models were evaluated under different setups. Both numbers are real; only one is comparable to anything.

Is Astra better than Claude Fable 5.1 at computer use?

Unknown from the published numbers. Astra's 72.6% is on OSWorld 2.0; Fable's 77.9% is on a different version of the test. The two cannot be subtracted.

What does Artificial Analysis measure differently?

They run their own harness across all models, which makes their comparisons internally consistent. On their index Astra scores 61 — level with GPT-5.6 Sol and behind Claude Fable 5.1 — with a large hallucination improvement and regressions on economic and support tasks.

Why does GDPval being missing matter?

It is OpenAI's own benchmark for real economic task value, and Astra is marketed on doing professional work. Its absence from launch materials means the claim closest to the product's positioning is the one with no published number behind it.

Which benchmark should I actually care about?

Whichever most resembles your workload — and then not much. The two most decision-relevant findings on this page came from independent testing rather than benchmarks: measured regressions on customer support, and detail absorption on long-running projects.

Is a 100% score good or suspicious?

Usually it means the benchmark is exhausted rather than the model is perfect. ExploitBench at 100% is less informative than ExploitGym at 42.4%, which is harder and still has headroom.

Work with me

I read the harness, not just the headline

Jayant Solanki

I'm Jayant Solanki — an SEO, GEO and automation strategist. Most of my work is separating what a system does from what it's marketed as doing, then building around the real mechanism.

Ranked #1 for "metal buildings" · +30% YoY organic traffic · Evidence-labelled research

Sources

Jayant Solanki

Jayant Solanki

AI-Ready SEO, GEO & AIO strategist based in Indore, India. Part of the GPT-6 Astra cluster — updated as independent replication lands.

More about Jayant →