# Gemini 4 Argon Benchmarks: Every Score and Who Ran It

> Canonical HTML version: https://thejayant.in/blog/gemini-4-argon-benchmarks
> Author: Jayant Solanki — https://thejayant.in/
> This Markdown file is a plain-text twin of the article at the URL above. Same content, no page furniture. It is public, not bot-only.

Confirmed Published by Google in its launch post or its evaluation methodology document.

Reported From named coverage or leaderboards I could not open myself.

Observed Measured by an independent party, such as Artificial Analysis.

Speculation My own reading of what the numbers mean.

Gemini 4 Argon leads **13 of the 19 benchmarks** Google published, ties one and trails on five. Its wins cluster in knowledge work, long context, science and video. Its losses cluster in **coding and computer use**: it comes last of four on FrontierSWE v2 (55.0%) and Terminal-Bench 4.0 (57.4%). Google ran several of Argon's own scores itself, rivals' numbers are mostly their own self-reported results, and GPT-6.1 Sol is not in the table. The one independent score so far, **53 on the Artificial Analysis Intelligence Index**, puts Argon 8th of 226 models, a strong model rather than a runaway leader.

This is the companion teardown to the [full Gemini 4 Argon guide](https://thejayant.in/blog/gemini-4-argon). The guide covers what the model is, who can use it and what it costs. This page covers what each number means, who measured it, and which comparisons are fair.

It needs its own page because Argon can't be independently tested yet: only vetted cyber defenders can use it. For now, almost every score comes from Google, so the way it was measured matters as much as the score itself.

## 1. The full Gemini 4 Argon benchmark table

These are the 19 results Google published on 30 September 2026, comparing Argon with OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 and Claude Opus 5.5. The last column says who produced the number. Confirmed

### Knowledge work

| Benchmark | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 | Measured by |
| --- | --- | --- | --- | --- | --- |
| Vals Index | **68.9%** | 63.1% | 65.8% | 67.0% | Vals AI |
| AutomationBench | **51.3%** | 41.4% | 31.4% | 42.5% | Zapier leaderboard |
| Vals Finance Agent v2 | **65.4%** | 53.5% | 58.9% | 58.6% | Vals AI |
| Harvey Legal Agent Benchmark | **19.6%** | 5.4% | 6.7% | 3.8% | Vals AI |

### Coding and ML engineering

| Benchmark | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 | Measured by |
| --- | --- | --- | --- | --- | --- |
| DeepSWE v1.1 | **77.9%** | 74.1% | 67.4% | 74.2% | Google ran Argon; rivals from leaderboard and system cards |
| FrontierSWE v2 | 55.0% | **65.5%** | 56.3% | 62.3% | Proximal leaderboard |
| Vibe Code Bench | **91.9%** | 89.6% | 90.3% | 90.3% | Vals AI |
| Terminal-Bench 4.0 | 57.4% | 58.2% | 57.9% | **66.4%** | Google ran Argon; rivals from leaderboard |
| PostTrainBench v1.1 | 45.3% | 44.3% | 40.2% | **49.3%** | Google ran all models |

### Science, maths and long context

| Benchmark | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 | Measured by |
| --- | --- | --- | --- | --- | --- |
| Terminal-Bench Science 0.1 | 57.6% | **68.1%** | 52.6% | 63.3% | Google ran Argon; rivals from leaderboard |
| LABBench 2 | **88.8%** | 85.4% | 68.6% | 73.1% | Google ran all models |
| RiemannBench | **76.0%** | 72.0% | 65.6% | 69.6% | Surge leaderboard |
| GraphWalks, up to 128K | **99.7%** | 98.7% | 91.4% | 90.6% | Google ran all models |
| GraphWalks, 256K to 1M | **84.2%** | 71.8% | 65.0% | 66.8% | Google ran all models |

### Computer use, multimodal and cybersecurity

| Benchmark | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 | Measured by |
| --- | --- | --- | --- | --- | --- |
| Agent's Last Exam | **39.5%** | 34.2% | — | 38.2% | Google ran Argon; rivals from leaderboard |
| OSWorld 2.0 (offline subset) | 69.2% | **72.6%** | — | — | Google ran Argon; Astra from OpenAI |
| Chartography | **71.6%** | 71.0% | 46.2% | 66.3% | Surge leaderboard |
| LVBench (long video) | **91.7%** | 87.5% | 79.7% | 83.7% | Google ran all models |
| CWE-bench v1 | **68%** | **68%** | 58% | 67% | CWE-bench leaderboard |

The scores are as published by Google; this table follows the reproduction by [Emergent](https://emergent.sh/learn/gemini-4-argon-benchmarks), checked against Google's launch post for every figure Google quotes in text. The "Measured by" column comes from Google's [evaluation methodology document](https://deepmind.google/models/evals-methodology/gemini-4-argon/). One coverage source lists GPT-6 Astra at 62.1% on Terminal-Bench 4.0 rather than 58.2%; either way, Argon is last on that row. Reported

## 2. Where Gemini 4 Argon wins

**Argon's clearest wins are in professional knowledge work, very long context and video.** These are also the rows where the margin is large enough to survive differences in test setup. Confirmed

- **Legal and finance.** Harvey's Legal Agent Benchmark is the most striking row: 19.6% against 3.8% to 6.7% for the rivals, so Argon solves about three times as many tasks. On Vals Finance Agent v2 it leads by 6.5 points. Both were run by Vals AI, not Google.
- **Very long inputs.** On GraphWalks between 256,000 and 1 million tokens of context, Argon scores 84.2% against 71.8% for the next best. Up to 128K, every model is above 90% and the test barely separates them. The long range is where the difference shows.
- **Long video.** 91.7% on LVBench, four points clear. Note that Google sampled video at 1 frame per second for Gemini but gave rivals a fixed number of frames (300 to 800) because of their API limits. That helps Gemini on long videos.
- **Business automation.** 51.3% on Zapier's AutomationBench, nearly nine points ahead, on Zapier's own private test set.

The pattern fits what Google built Argon for: long, careful tasks where holding a lot of material in mind matters more than speed. Speculation

## 3. Where Gemini 4 Argon loses

**Argon is not the best coding model.** On the two hardest agentic coding tests in Google's own table, it comes last of four. Confirmed

- **FrontierSWE v2: 55.0%**, against 65.5% for GPT-6 Astra and 62.3% for Opus 5.5. This one comes from Proximal's public leaderboard, not from Google.
- **Terminal-Bench 4.0: 57.4%**, nine points behind Opus 5.5. Google ran Argon's score itself.
- **Terminal-Bench Science 0.1: 57.6%**, ten points behind Astra, despite Google using a six-times-longer verifier timeout for Argon.
- **OSWorld 2.0: 69.2%**, behind Astra's 72.6%. Google reports a partial score on the offline subset, the best of three runs, so it is not directly comparable to Anthropic's combined figure.
- **PostTrainBench: 45.3%**, behind Opus 5.5 on a test of improving another model, which Google ran for all four models.

Coverage of the launch noted that Argon trailed on two of the four coding benchmarks Google chose to include, and Reuters reported Google described Argon as larger than its earlier Pro models. Reported

DeepSWE v1.1, Argon's best coding row at 77.9%, deserves a second look. Google computed Argon's score with a minimal agent harness, while rivals' scores come from a public leaderboard and system cards. The win is real on Google's numbers, but it is the coding win with the most setup difference behind it. Speculation

## 4. Who ran each number

**Google's evaluation document is unusually clear about its own method, and it shows that the table mixes three kinds of measurement.** Confirmed

| How it was measured | Rows | How much to trust the comparison |
| --- | --- | --- |
| Third party ran it for every model | Vals Index, Finance Agent, Harvey Legal, Vibe Code Bench, AutomationBench, FrontierSWE v2, RiemannBench, Chartography, CWE-bench v1 | Highest: same referee, though harnesses can still differ |
| Google ran every model itself | PostTrainBench, LABBench 2, GraphWalks, LVBench | Medium: same setup, but chosen by Google |
| Google ran Argon; rivals from their own reports or leaderboards | DeepSWE, Terminal-Bench 4.0, Terminal-Bench Science, Agent's Last Exam, OSWorld 2.0 | Lowest: different people, different runs |

Three method details matter when reading the table:

- **Highest thinking settings.** Argon was run at its highest thinking setting through the Gemini API. Rivals were reported at their maximum reasoning setting where available, otherwise their best available result.
- **Pass@1, single attempt.** No majority voting or parallel attempts, except where noted. Smaller benchmarks were averaged over several trials.
- **Safety filters on.** In computer-use tests, responses blocked by Argon's safety filters counted as empty answers. That can only lower Argon's score.

None of this is unusual. Every lab publishes launch numbers this way. But it means the third-party rows are the ones to quote, and the rows where Google ran only Argon are the ones to treat as provisional. Speculation

## 5. The independent read so far

**The only independent measurement I could verify directly is Artificial Analysis, which scores Gemini 4 Argon at 53 on its Intelligence Index, 8th of 226 models.** It calls the result well above average, against a median of 26. Observed ([Artificial Analysis](https://artificialanalysis.ai/models/gemini-4-argon))

- **Cost to run its index:** about $1.99 per task at the $2 / $10 introductory price.
- **Verbosity:** 110 million output tokens to complete the index, against a median of 81 million. Argon thinks at length, which costs money even at a low token price. Observed
- **Against rivals:** coverage of the same index puts GPT-6 Astra and Claude Fable 5.1 at the same score, 53, with Claude Opus 5.5 ahead at 58. Reported
- **Coding arenas:** one analysis, citing The Decoder, puts Argon 8th on Code Arena's WebDev board, behind GPT-6.1 Sol. Reported

Independent results agree with Google's own table on the big picture: Argon is a frontier-level model that leads in some areas and is not the strongest coder. The gap on the Intelligence Index is a reminder that Google's 13-of-19 headline reflects which benchmarks were chosen. Speculation

## 6. What is missing from the comparison

- **GPT-6.1 Sol.** OpenAI's newest model, launched the day before Argon, is not in Google's table. Its own reported DeepSWE v1.1 score is around 75%, close to Argon's 77.9%, and it is available today at the same $2 / $10 price. See the [DevDay explainer](https://thejayant.in/blog/openai-devday-2026-announcements#sol).
- **Claude Sonnet 5.5.** Launched two days before Argon, and an independent evaluator has reported it ahead of Argon on app building and code migration. Reported
- **Common reasoning tests.** There is no Humanity's Last Exam, GPQA or ARC-AGI score for Argon in Google's table.
- **Independent replication.** With access limited to Fairwind members, nobody outside Google can rerun the self-computed rows yet.
- **Cost per task by benchmark.** Google gives scores but not how many tokens or how much time each took. With Argon's verbosity, that is the number buyers need most.

## 7. Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5, by job

**On Google's published numbers, Argon is the strongest of the three for research-heavy knowledge work and long inputs, Opus 5.5 is the strongest terminal coder, and GPT-6 Astra is strongest on computer use and FrontierSWE.**

| Job | Leader in Google's table | Deciding rows |
| --- | --- | --- |
| Legal and finance research | Gemini 4 Argon | Harvey Legal, Finance Agent v2, Vals Index |
| Business process automation | Gemini 4 Argon | AutomationBench |
| Very long documents | Gemini 4 Argon | GraphWalks 256K–1M |
| Terminal and repository coding | Claude Opus 5.5 | Terminal-Bench 4.0, PostTrainBench |
| Hard software engineering tasks | GPT-6 Astra | FrontierSWE v2 |
| Operating a computer | GPT-6 Astra | OSWorld 2.0 |
| Video and charts | Gemini 4 Argon | LVBench, Chartography |
| Vulnerability patching | Tie: Argon and Astra | CWE-bench v1 |

Remember the access gap when you read this: Astra and Opus 5.5 are available today, Argon is not. For a practical recommendation, see [which model to use now](https://thejayant.in/blog/gemini-4-argon#which) in the main guide, and for the earlier generation, the [GPT-6 Astra vs Claude Fable vs Gemini comparison](https://thejayant.in/blog/gpt-6-astra-vs-claude-fable-vs-gemini).

## 8. How to read any frontier model benchmark

The same five questions work for Argon, Astra, Opus or the next launch:

1. **Who ran it?** A third-party leaderboard beats a lab's own run of its own model.
2. **Was every model run the same way?** Same harness, same thinking setting, same number of attempts. If not, the comparison is approximate.
3. **Is the benchmark near its ceiling?** When every model scores above 90%, as on GraphWalks up to 128K, the row tells you very little.
4. **Which models are missing?** A table without the newest rival is a table chosen at a convenient moment.
5. **What did it cost?** A higher score that takes twice the tokens may be the worse deal. Ask for cost per finished task.

The [GPT-6 Astra benchmark teardown](https://thejayant.in/blog/gpt-6-astra-benchmarks) applies the same checklist to OpenAI's numbers, including a 37-point swing caused by the test harness alone.

## 9. Frequently asked questions

### How does Gemini 4 Argon score on benchmarks?

In the 19 benchmarks Google published on 30 September 2026, Gemini 4 Argon leads 13, ties 1 and trails on 5. It scores 68.9% on the Vals Index, 77.9% on DeepSWE v1.1, 84.2% on GraphWalks 256K–1M and 91.7% on LVBench. It trails on FrontierSWE v2 (55.0%), Terminal-Bench 4.0 (57.4%), Terminal-Bench Science 0.1, PostTrainBench and OSWorld 2.0.

### Is Gemini 4 Argon good at coding?

Gemini 4 Argon is strong at coding but not the leader. It tops DeepSWE v1.1 at 77.9% and Vibe Code Bench at 91.9%, but comes last of four on FrontierSWE v2 at 55.0% and on Terminal-Bench 4.0 at 57.4%, where Claude Opus 5.5 leads at 66.4%.

### What is Gemini 4 Argon's Artificial Analysis Intelligence Index score?

Gemini 4 Argon scores 53 on the Artificial Analysis Intelligence Index, ranking 8th of 226 models. Artificial Analysis notes it is "somewhat verbose", using about 110 million output tokens to run the index against a median of 81 million, at a cost of about $1.99 per task.

### Are Gemini 4 Argon's benchmarks independently verified?

Partly. Nine of the 19 rows come from third-party leaderboards such as Vals AI, Zapier, Proximal and Surge. Google ran the rest itself, either for all models or only for Argon. Because Argon is limited to Fairwind Program members, outside groups cannot yet rerun Google's own measurements.

### Is Gemini 4 Argon better than Claude Opus 5.5?

On Google's table, Gemini 4 Argon beats Claude Opus 5.5 on 14 of the 18 rows where both are listed, including legal, finance and long-context work. Opus 5.5 leads on FrontierSWE v2, Terminal-Bench 4.0, Terminal-Bench Science 0.1 and PostTrainBench. On the independent Artificial Analysis Intelligence Index, Opus 5.5 scores higher.

### Why is GPT-6.1 Sol not in Google's comparison?

Google compared Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. GPT-6.1 Sol launched on 29 September 2026, one day before Argon, and Google did not include it. GPT-6.1 Sol's own reported DeepSWE v1.1 score is around 75%, close to Argon's 77.9%.

Quote Argon's third-party rows, treat the rows Google ran itself as provisional, and judge cost per finished task, not the headline 13-of-19.

## Sources

**Editorial note:** all scores are as published on 30 September 2026 and reproduced from Google's materials and named secondary sources. I have not run any benchmark myself. This page will be updated as independent results arrive.
