# GPT-6 Astra Benchmarks Explained: Every Score, and Why the Records Have Asterisks

> Canonical HTML version: https://thejayant.in/blog/gpt-6-astra-benchmarks
> Author: Jayant Solanki — https://thejayant.in/
> This Markdown file is a plain-text twin of the article at the URL above. Same content, no page furniture. It is public, not bot-only.

GPT-6 Astra sets records on the things it was built for — computer use, terminals, CAD, cyber — and is statistically flat on general intelligence. The headline ARC-AGI-3 number is 99.9% under OpenAI's own Provider Adapter harness and **62.7% on the standard one**, and rival models were scored under different setups. The OSWorld comparison against Claude Fable 5.1 uses a different version of the test. GDPval is missing entirely. None of it has been independently replicated. It is a genuinely strong specialised model with an unusually noisy scoreboard.

This is the companion teardown to the [full GPT-6 Astra guide](https://thejayant.in/blog/gpt-6-astra). The pillar covers what the model is; this page covers what the numbers mean, which ones are comparable, and which ones quietly aren't.

The reason it needs its own page: most launch coverage reprinted the scores in a table and moved on. The scores are fine. The _harnesses_ are the story.

## What's covered

1. [The full benchmark table](#full)
2. [The ARC-AGI-3 problem](#arc)
3. [The OSWorld version mismatch](#osworld)
4. [Artificial Analysis: the independent read](#aa)
5. [What's missing](#missing)
6. [Head-to-head against the field](#head)
7. [How to read any frontier benchmark](#checklist)
8. [FAQ](#faq)

## 1. The full benchmark table

Everything OpenAI published, grouped by family, with comparisons where they exist and are meaningful.

### Computer use — the category Astra was built to win

| Benchmark | Astra | GPT-5.6 Sol | Notes |
| --- | --- | --- | --- |
| OSWorld 2.0 | 72.6% | 65.7% | ~40 min/task vs ~75 min — see §3 |
| Agents' Last Exam | 59.3% | — | Claude Opus 5: 55.5% |
| ScreenSpot-Pro | 92.7% | — | UI element grounding |

The time-per-task figure matters more than the accuracy figure here. A 6.9-point accuracy gain is respectable; doing it in 47% less wall-clock time is what changes whether an agent workflow is economically viable.

### Coding and engineering

| Benchmark | Astra | Comparison |
| --- | --- | --- |
| Terminal-Bench 4.0 | 57.9% | Sol: 37.3% |
| DeepSWE v1.1 | 74.1% | Sol: 70.8%; Meta Muse 1.3 reportedly 75.4% |
| FrontierCode Extended | 64.5% | — |

Terminal-Bench is the standout — a 20-point jump is not incremental. Note also that on DeepSWE, the benchmark OpenAI leans on hardest for its cost-per-task argument, Astra is **behind** a reported Meta Muse score. That comparison did not feature prominently in the launch materials.

### Science and reasoning

| Benchmark | Astra | Comparison |
| --- | --- | --- |
| FrontierMath Tier 4 v2 | 97.6% | — |
| GPQA Diamond | 96.0% | — |
| ARC-AGI-3 | 99.9% / 62.7% | Depends entirely on harness — see §2 |
| Terminal-Bench Science 0.1 | 64.6% | Claude Fable 5.1: 52.6% |
| BenchCAD Vision2Code | 95.9% | Fable 5.1: 84.3% |
| Humanity's Last Exam | 57.2% | — |

GPQA Diamond at 96% is effectively saturated — when a benchmark approaches its ceiling it stops discriminating between models, and should be retired from comparison rather than quoted as a win.

### Cybersecurity

| Benchmark | Astra | GPT-5.6 Sol |
| --- | --- | --- |
| ExploitBench | 100% | 78.5% |
| ExploitGym | 42.4% | 30.3% |
| SRE-Bench (single attempt) | 88.0% | — |

A 100% score should always prompt the same question: is the benchmark exhausted, or is the model that good? ExploitGym at 42.4% — harder, more open-ended — is the more informative number, and it is the one that justifies the Critical classification alongside the two zero-days found during evaluation.

### Alignment and safety

| Metric | Astra | Prior |
| --- | --- | --- |
| Safety benchmark error rate | 2.4% | — |
| Circumvention attempts | 0% | Sol: 48.2–56% |
| Hallucination benchmark error rate | 4.2% | — |
| Cyber jailbreak rejection | 91.5% | 59% |

## 2. The ARC-AGI-3 problem

This is the number that made the headlines and the one that survives scrutiny least well.

| Provider Adapter harness | 99.9% |
| --- | --- |
| Standard harness | 62.7% |
| Gap | 37.2 points, same model |

OpenAI ran Astra through its own Responses API harness with two settings changed "to better reflect how the model performs in real-world use." The Provider Adapter, as Simon Willison flagged, "preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work."

Three things follow:

1. **It is a real engineering result.** Persisting reasoning state across requests genuinely helps, and it reflects how you would actually deploy the model. Dismissing the number entirely is as wrong as quoting it uncritically.
2. **It is not a like-for-like comparison.** The rival models in the same chart ran under different setups. A 37-point swing from harness configuration is larger than the gap between most frontier models.
3. **Reported figures vary.** Some outlets cite 98.6% rather than 99.9%, likely different runs or settings. The variance itself is informative.

ARC Prize was unusually direct about the interpretation: _"we are not claiming that it is AGI,"_ noting the benchmark has "deterministic, closed-ended mechanics" that don't represent real-world complexity. Greg Kamradt's more interesting finding got far less coverage: on ARC-AGI-3, Astra **surpassed the human action-efficiency baseline on 96% of levels** — a claim about solving unfamiliar environments in fewer moves than people, which is a different and arguably more impressive thing than raw accuracy.

**A benchmark score without its harness is not a number, it's a marketing asset.** From now on, any frontier score you cite should carry the harness, the settings and whether competitors ran the same configuration. If a vendor won't publish that, treat the score as unverified.

## 3. The OSWorld version mismatch

The second comparability problem, and a quieter one.

Astra scores **72.6% on OSWorld 2.0**. Anthropic reported **77.9% for Claude Fable 5.1 on OSWorld** — a higher number, on a different version of the test.

Read carelessly, that says Fable beats Astra at computer use. Read carefully, it says nothing at all: benchmark versions change task sets, scoring and difficulty. You cannot subtract one from the other.

What you _can_ say with confidence is the internally consistent comparison: Astra 72.6% vs GPT-5.6 Sol 65.7% on the same version, at roughly half the time per task. That is a real generational gain within one vendor's own lineage — and it is the claim OpenAI is entitled to make.

## 4. Artificial Analysis: the independent read

The most useful third-party evidence available, because they run their own harness across every model — which is exactly the property the vendor numbers lack.

| Measure | Result | What it means |
| --- | --- | --- |
| Intelligence Index | 61 | Level with GPT-5.6 Sol; behind Claude Fable 5.1 |
| Coding | 67 | Matching Claude Opus 5 and Fable 5 |
| Token efficiency | ~1/3 of Sol's tokens | In the Codex harness, at roughly equal cost, for better results |
| Hallucination rate | 92% → 51% at max effort | Accuracy up 4 points alongside it |
| Economic / support tasks | Regressions | Measurably worse than the predecessor |

Two of those deserve more attention than they got.

**The hallucination collapse is the most underrated number in the launch.** A drop from 92% to 51% on their measure, with accuracy improving simultaneously, is a larger practical improvement for most real workloads than any computer-use record. Fewer confident wrong answers changes how much review a human has to do, which is the actual cost driver in production.

**The regressions matter more than the records for most buyers.** If your workload is customer support, this model is measurably worse than its predecessor at 2.5× the price. That is the single most actionable finding on this page.

Their conclusion — Astra is a **specialised** model, not a universally superior one — is the honest summary of the whole scoreboard.

## 5. What's missing

- **GDPval.** OpenAI's own benchmark for real economic task value is absent from the launch materials. For a model sold on doing professional work, that is conspicuous. Absence of a result is not evidence of a bad result — but it is not nothing either.
- **Parameter count, training data size, architecture.** "Opaque recurrence" has a name and no technical documentation.
- **Task-level cost data.** Without it, OpenAI's cost-per-task savings claims cannot be verified by anyone outside OpenAI.
- **Independent replication.** As of publication, no third party outside Artificial Analysis's separate index has reproduced these scores.
- **Long-running production behaviour** with interruptions, ambiguous permissions and repeated tool failures. Every published benchmark is a clean run.

## 6. Head-to-head against the field

|  | GPT-6 Astra | Claude Fable 5.1 | GPT-5.6 Sol | Meta Muse |
| --- | --- | --- | --- | --- |
| Price per 1M in/out | $10 / $50 | $10 / $50 | $4 / $20 | $1.25 / $4.25 (Spark) |
| AA Intelligence Index | 61 | Ahead of Astra | ~61 | — |
| Computer use | 72.6% OSWorld 2.0 | 77.9% on a different version | 65.7% | — |
| DeepSWE v1.1 | 74.1% | — | 70.8% | 75.4% reported (1.3) |
| Science terminal | 64.6% | 52.6% | — | — |
| Vision-to-code (BenchCAD) | 95.9% | 84.3% | — | — |

The summary that survives all the caveats: **Astra and Fable 5.1 cost exactly the same and win different things.** Astra takes CAD, science terminals and cyber. Fable takes general intelligence. Sol remains 60% cheaper and is not obviously worse for chat and general reasoning. Muse is an order of magnitude cheaper and competitive on the one coding benchmark where a comparison exists.

That is not a leaderboard. It is a routing table — and routing is what you should be building, rather than picking a winner.

## 7. How to read any frontier benchmark from now on

### The six questions

- **What harness produced this number?** If it's the vendor's own, was it the same one used for the competitors in the chart?
- **What version of the benchmark?** OSWorld 2.0 and OSWorld are not the same test.
- **Is the benchmark saturated?** Anything above ~95% has stopped discriminating.
- **Which results are absent?** A missing benchmark from a vendor that owns it is a data point.
- **Has anyone independent reproduced it?** Vendor numbers are a hypothesis until then.
- **Was failure included?** Clean runs tell you nothing about behaviour under tool failures, interruptions or ambiguous permissions.

Apply those six to any launch and most scoreboards get considerably quieter. That is the correct outcome.

## 8. Frequently asked questions

**Did GPT-6 Astra really score 99.9% on ARC-AGI-3?**

Under OpenAI's own Provider Adapter harness, yes. On the standard harness the same model scores 62.7%, and competing models were evaluated under different setups. Both numbers are real; only one is comparable to anything.

**Is Astra better than Claude Fable 5.1 at computer use?**

Unknown from the published numbers. Astra's 72.6% is on OSWorld 2.0; Fable's 77.9% is on a different version of the test. The two cannot be subtracted.

**What does Artificial Analysis measure differently?**

They run their own harness across all models, which makes their comparisons internally consistent. On their index Astra scores 61 — level with GPT-5.6 Sol and behind Claude Fable 5.1 — with a large hallucination improvement and regressions on economic and support tasks.

**Why does GDPval being missing matter?**

It is OpenAI's own benchmark for real economic task value, and Astra is marketed on doing professional work. Its absence from launch materials means the claim closest to the product's positioning is the one with no published number behind it.

**Which benchmark should I actually care about?**

Whichever most resembles your workload — and then not much. The two most decision-relevant findings on this page came from independent testing rather than benchmarks: measured regressions on customer support, and detail absorption on long-running projects.

**Is a 100% score good or suspicious?**

Usually it means the benchmark is exhausted rather than the model is perfect. ExploitBench at 100% is less informative than ExploitGym at 42.4%, which is harder and still has headroom.

## Sources
