October 6, 2026

Gemini 4 Argon Benchmarks Explained: Coding, Finance, Legal, Science & Cybersecurity

Gemini 4 Argon Benchmarks Explained: Coding, Finance, Legal, Science & Cybersecurity

Gemini 4 Argon Benchmarks Explained: Coding, Finance, Legal, Science & Cybersecurity

Google’s Gemini 4 Argon has arrived with one of the strongest benchmark packages the company has published for a frontier model.

According to Google DeepMind’s official evaluation table, Argon leads or ties the top score across many of the tests it published against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5.

Its strongest areas include:

  • enterprise knowledge work;
  • financial research;
  • legal agent tasks;
  • long-horizon software engineering;
  • long-context reasoning;
  • multimodal understanding; and
  • defensive cybersecurity.

But the results are not a clean sweep.

GPT-6 Astra beats Argon on some difficult software-engineering, scientific terminal and computer-use evaluations, while Claude Opus 5.5 leads on Terminal-Bench 4.0 and PostTrainBench in Google’s comparison.

Independent Artificial Analysis also now places Gemini 4 Argon among the frontier leaders, giving the model an Intelligence Index score of 53 at its High setting.

The important question is therefore not:

“Does Argon win benchmarks?”

It clearly wins several.

The more useful question is:

“Which benchmarks matter for the work you actually want the model to do?”

Gemini 4 Argon Benchmark Results at a Glance

BenchmarkCategoryGemini 4 ArgonGPT-6 AstraClaude Opus 5.5
Vals IndexKnowledge work68.9%63.1%67.0%
AutomationBenchKnowledge work51.3%41.4%42.5%
Vals Finance Agent v2Finance65.4%53.5%58.6%
Harvey Legal Agent BenchmarkLegal19.6%5.4%3.8%
DeepSWE v1.1Agentic coding77.9%74.1%74.2%
FrontierSWE v2Agentic coding55.0%65.5%62.3%
Vibe Code BenchAgentic coding91.9%89.6%90.3%
Terminal-Bench 4.0Agentic coding57.4%58.2%66.4%
PostTrainBenchML engineering45.3%44.3%49.3%
Terminal-Bench Science 0.1Science57.6%68.1%63.3%
LABBench 2Science & math88.8%85.4%73.1%
RiemannBenchScience & math76.0%72.0%69.6%
GraphWalks up to 128KLong context99.7%98.7%90.6%
GraphWalks 256K–1MLong context84.2%71.8%66.8%
Agent’s Last ExamComputer use39.5%34.2%38.2%
OSWorld-2.0 offline subsetComputer use69.2%72.6%—
ChartographyMultimodal71.6%71.0%66.3%
LVBenchMultimodal91.7%87.5%83.7%
CWE-bench v1Cybersecurity68.0%68.0%67.0%

These are the scores published on the official Google DeepMind Gemini 4 Argon model page.

Google also publishes a separate Gemini 4 Argon evaluation methodology document explaining how the scores were produced or sourced.

Important: Not Every Score Was Independently Produced by Google

Before interpreting the table, it is important to understand Google’s methodology.

Google says most Gemini 4 Argon scores were run using the Gemini API at the highest thinking settings and are generally reported as pass@1 unless noted otherwise.

However, the comparison table combines several types of evidence.

Depending on the benchmark, scores can come from:

  • Google’s own internal runs;
  • official public benchmark leaderboards;
  • Vals AI;
  • Zapier;
  • Surge;
  • other model providers’ published results; and
  • benchmark-specific official sources.

For example, Google says its DeepSWE v1.1 result for Argon was self-computed, while competing model scores were taken from the relevant public leaderboard or system cards.

Google also says GraphWalks scores for all models were self-computed under its evaluation setup.

This does not make the results invalid, but it means readers should avoid treating the table as though one independent laboratory ran every model under perfectly identical conditions.

Gemini 4 Argon Coding Benchmarks

Coding is one of Argon’s strongest categories, but the results reveal an important pattern.

Argon is extremely strong at long-horizon software engineering and code-generation tasks, but it does not lead every terminal-driven engineering benchmark.

DeepSWE v1.1: Argon 77.9%

On DeepSWE v1.1, Google reports:

Gemini 4 Argon: 77.9%

Claude Opus 5.5: 74.2%

GPT-6 Astra: 74.1%

Claude Fable 5.1: 67.4%

DeepSWE focuses on real-world, long-horizon software-engineering work.

Argon’s 77.9% result gives it a:

3.8 percentage-point lead over GPT-6 Astra

and a:

3.7-point lead over Claude Opus 5.5

This supports Google’s positioning of Argon as a strong model for repository-level work, multi-file changes and long-running software-engineering tasks.

Vibe Code Bench: Argon 91.9%

Google reports:

Gemini 4 Argon: 91.9%

Claude Opus 5.5: 90.3%

Claude Fable 5.1: 90.3%

GPT-6 Astra: 89.6%

The differences here are much smaller.

Argon’s advantage over Astra is:

91.9 − 89.6 = 2.3 percentage points

That is a useful signal, but not enough by itself to establish that Argon will outperform Astra on every coding task.

Where Gemini 4 Argon Loses on Coding

The most important coding loss is FrontierSWE v2.

FrontierSWE v2

Google’s table shows:

GPT-6 Astra: 65.5%

Claude Opus 5.5: 62.3%

Claude Fable 5.1: 56.3%

Gemini 4 Argon: 55.0%

That means Astra leads Argon by:

65.5 − 55.0 = 10.5 percentage points

This is a large gap and a reminder that “best coding model” depends heavily on the benchmark and harness.

Terminal-Bench 4.0

Google reports:

Claude Opus 5.5: 66.4%

GPT-6 Astra: 58.2%

Claude Fable 5.1: 57.9%

Gemini 4 Argon: 57.4%

Argon does not lead terminal-based agentic coding in this table.

This suggests that developers should benchmark Argon separately for repository planning and terminal-heavy execution rather than assuming one coding score represents both.

For more on Astra’s coding and tool capabilities, see our GPT-6 Astra complete guide.

Finance Benchmarks: One of Argon’s Strongest Areas

Google is positioning Gemini 4 Argon heavily toward professional financial research.

On Vals Finance Agent v2, the published scores are:

Gemini 4 Argon: 65.4%

Claude Fable 5.1: 58.9%

Claude Opus 5.5: 58.6%

GPT-6 Astra: 53.5%

Argon’s lead over Astra is:

65.4 − 53.5 = 11.9 percentage points

That is one of Argon’s more meaningful published enterprise advantages.

Vals Finance Agent is designed around multi-step financial research rather than simple question answering.

That makes the result especially relevant for potential workflows such as:

  • company research;
  • financial document analysis;
  • earnings research;
  • market research;
  • investment-banking workflows; and
  • multi-source financial analysis.

OpenAI is also targeting financial workflows with Astra. We cover that separately in our ChatGPT for Financial Services guide.

The most dramatic score difference in Google’s table comes from Harvey’s Legal Agent Benchmark.

Google reports:

Gemini 4 Argon: 19.6%

Claude Fable 5.1: 6.7%

GPT-6 Astra: 5.4%

Claude Opus 5.5: 3.8%

Relative to Astra, Argon’s score is more than three times as high.

But this benchmark needs careful interpretation.

A score of 19.6% still means the model does not successfully complete the majority of benchmark tasks.

So the correct conclusion is not:

“Gemini 4 can replace lawyers.”

A more defensible conclusion is:

“Argon shows a large relative lead on this specific legal-agent benchmark, while absolute performance still leaves substantial room for improvement.”

High-stakes legal work should continue to involve qualified human review.

General Knowledge Work: Vals Index and AutomationBench

Google reports Argon at:

68.9% on the Vals Index

compared with:

67.0% for Claude Opus 5.5

65.8% for Claude Fable 5.1

63.1% for GPT-6 Astra

The Vals Index is intended to measure economically relevant professional work.

Argon also leads Google’s AutomationBench comparison:

Argon: 51.3%

Opus 5.5: 42.5%

Astra: 41.4%

Fable 5.1: 31.4%

These results support Google’s claim that Argon is designed less as a consumer chatbot and more as a professional workflow model.

Long-Context Benchmarks: Argon’s Largest Technical Lead

Long-context performance is one of Argon’s strongest published areas.

Google uses GraphWalks, which requires models to navigate graph structures described inside long prompts.

GraphWalks up to 128K

Google reports:

Argon: 99.7%

Astra: 98.7%

Fable 5.1: 91.4%

Opus 5.5: 90.6%

The frontier models are relatively close in this shorter subset.

GraphWalks from 256K to 1M tokens

The gap widens sharply on the longer-context subset:

Argon: 84.2%

Astra: 71.8%

Opus 5.5: 66.8%

Fable 5.1: 65.0%

Argon’s lead over Astra is:

84.2 − 71.8 = 12.4 percentage points

This is one of the strongest reasons to pay attention to Argon for:

  • large code repositories;
  • very long legal records;
  • enterprise knowledge bases;
  • large document sets; and
  • long-running research workflows.

Science and Math: Strong, but Not a Sweep

Argon leads Google’s comparison on LABBench 2 and RiemannBench.

LABBench 2

Argon: 88.8%

Astra: 85.4%

Opus 5.5: 73.1%

Fable 5.1: 68.6%

RiemannBench

Argon: 76.0%

Astra: 72.0%

Opus 5.5: 69.6%

Fable 5.1: 65.6%

However, Astra wins Terminal-Bench Science 0.1:

Astra: 68.1%

Opus 5.5: 63.3%

Argon: 57.6%

Fable 5.1: 52.6%

Astra’s lead over Argon is:

10.5 percentage points

This suggests Argon may be particularly strong at science reasoning, while Astra may have an advantage in scientific tasks that require active terminal execution.

Cybersecurity Benchmark: Argon and Astra Tie

Google has emphasized defensive cybersecurity more than almost any other Argon capability.

On CWE-bench v1, however, the published result is essentially a tie at the top:

Gemini 4 Argon: 68%

GPT-6 Astra: 68%

Claude Opus 5.5: 67%

Claude Fable 5.1: 58%

CWE-bench measures vulnerability remediation.

Google also says Argon performed strongly on internal vulnerability-discovery tests and on a black-box penetration-testing benchmark developed with Wiz, but those internal evaluations are not directly comparable to a fully public leaderboard.

The public CWE-bench result is therefore the cleaner comparison for outside readers.

Multimodal Benchmarks: Argon Leads Google’s Table

Argon also posts strong multimodal results.

LVBench

Argon: 91.7%

Astra: 87.5%

Opus 5.5: 83.7%

Fable 5.1: 79.7%

Chartography

Argon: 71.6%

Astra: 71.0%

Opus 5.5: 66.3%

Fable 5.1: 46.2%

Argon’s LVBench lead is more meaningful than the extremely narrow 0.6-point Chartography advantage over Astra.

Computer Use: Mixed Results

On Agent’s Last Exam, Google reports:

Argon: 39.5%

Opus 5.5: 38.2%

Astra: 34.2%

Argon leads this particular test.

But on Google’s OSWorld-2.0 offline subset comparison:

Astra: 72.6%

Argon: 69.2%

So computer-use performance is not uniformly in Argon’s favor.

For production browser and computer workflows today, Astra also has the practical advantage of wider availability and mature public tool support.

See our GPT-6 Astra for Work guide for more on computer use and professional workflows.

Independent Results: Artificial Analysis Scores Argon at 53

Google’s published benchmarks are useful, but independent evaluations matter because they reduce reliance on one model provider’s chosen tests.

Artificial Analysis currently gives Gemini 4 Argon at its High reasoning setting an:

Artificial Analysis Intelligence Index score of 53

That places it among the leading frontier models in the independent evaluator’s current ranking.

Artificial Analysis currently reports the following Argon results:

Independent testGemini 4 Argon High
Artificial Analysis Intelligence Index53
AutomationBench-AA78%
Terminal-Bench 4.057%
SciCode62%
Humanity’s Last Exam57%
GDP.pdf22%
AA-LCR long-context reasoning80%

Artificial Analysis also estimates a cost of approximately $1.99 per Intelligence Index task at Argon’s current discounted token pricing.

Independent source: Artificial Analysis — Gemini 4 Argon.

What the Independent Scores Change

The independent results strengthen the case that Argon is genuinely a frontier-tier model rather than simply looking strong in Google’s own launch materials.

However, they do not independently reproduce every Google-published benchmark.

For example, Artificial Analysis uses its own evaluation suite, reasoning settings and methodology.

So a responsible interpretation is:

Google’s benchmark table shows Argon’s claimed strengths by workload, while independent Artificial Analysis confirms that Argon broadly belongs near the top of the current model landscape.

Why Benchmark Results Can Disagree

AI benchmarks are not interchangeable.

A model can lead one coding benchmark and lose another because the tests measure different skills.

Differences can come from:

  • agent harnesses;
  • reasoning settings;
  • tool access;
  • time budgets;
  • number of attempts;
  • prompt construction;
  • verification methods;
  • benchmark contamination; and
  • how much compute is allowed during inference.

Google’s methodology document explicitly notes that some evaluations were self-computed while others use public leaderboards or provider-reported results.

This is why one benchmark should never be treated as a complete measurement of a model’s intelligence.

Which Gemini 4 Argon Benchmarks Matter Most?

That depends on your use case.

For software engineers

Pay the most attention to:

  • DeepSWE v1.1;
  • FrontierSWE v2;
  • Terminal-Bench 4.0; and
  • Vibe Code Bench.

The mixed results suggest you should test your own codebase rather than rely on one score.

For finance teams

Vals Finance Agent v2 is one of the most directly relevant published tests.

Argon’s 65.4% score is significantly higher than the competing scores in Google’s table.

Harvey’s Legal Agent Benchmark is relevant, but absolute scores remain low.

Use the result as a relative capability signal, not evidence that human legal review can be removed.

For enterprise knowledge work

The Vals Index, AutomationBench and long-context GraphWalks results are especially relevant.

For cybersecurity teams

CWE-bench v1 is the clearest public comparison, where Argon ties Astra at 68%.

For scientific workflows

Look at both reasoning-heavy evaluations such as LABBench 2 and execution-oriented tests such as Terminal-Bench Science.

The different winners show why benchmark selection matters.

Does Gemini 4 Argon Beat GPT-6 Astra?

On Google’s published benchmark table, Argon leads Astra across many categories, particularly:

  • knowledge work;
  • finance;
  • legal workflows;
  • DeepSWE;
  • long-context GraphWalks;
  • multimodal understanding; and
  • Agent’s Last Exam.

Astra leads Argon on important tests including:

  • FrontierSWE v2;
  • Terminal-Bench Science 0.1;
  • Terminal-Bench 4.0 by a narrow margin; and
  • OSWorld-2.0 offline subset.

The fairest conclusion is therefore:

Gemini 4 Argon appears stronger across Google’s enterprise and long-context benchmark set, while GPT-6 Astra retains advantages in several execution-heavy software, science and computer-use tests.

For our broader model comparison, see the Elite Era Trends GPT-6 Astra guide.

Should You Trust Gemini 4 Argon’s Benchmark Scores?

You should take them seriously, but not uncritically.

Three things are true at the same time:

1. Google published unusually detailed benchmark and methodology information.

2. Some results were self-computed or combine scores sourced under different evaluation setups.

3. Independent Artificial Analysis now confirms that Argon is broadly a frontier-tier model.

The best practice is therefore to use public benchmarks for model selection, then run a private evaluation on your own tasks before committing substantial production traffic.

Frequently Asked Questions

What is Gemini 4 Argon’s DeepSWE score?

Google reports a 77.9% score on DeepSWE v1.1, ahead of GPT-6 Astra at 74.1% and Claude Opus 5.5 at 74.2% in Google’s comparison.

What is Gemini 4 Argon’s finance benchmark score?

Google reports 65.4% on Vals Finance Agent v2, compared with 53.5% for GPT-6 Astra and 58.6% for Claude Opus 5.5.

Google reports 19.6% on Harvey’s Legal Agent Benchmark. That is much higher than the competing scores in Google’s table, although the absolute result remains below 20%.

What is Gemini 4 Argon’s cybersecurity score?

Google reports 68% on CWE-bench v1, tying GPT-6 Astra and slightly ahead of Claude Opus 5.5 at 67%.

What is Gemini 4 Argon’s long-context score?

Google reports 84.2% on GraphWalks for the 256K-to-1M-token subset, compared with 71.8% for GPT-6 Astra.

Does Gemini 4 Argon beat GPT-6 Astra at coding?

It depends on the benchmark. Argon leads DeepSWE v1.1 and Vibe Code Bench, while Astra leads FrontierSWE v2 and narrowly leads Terminal-Bench 4.0 in Google’s published comparison.

Does Gemini 4 Argon beat Claude Opus 5.5?

Argon leads Opus on several knowledge-work, long-context and multimodal benchmarks in Google’s table, while Opus leads Argon on Terminal-Bench 4.0 and PostTrainBench.

Are Google’s Gemini 4 Argon benchmarks independently verified?

Not all of them. Google’s methodology uses a mix of self-computed results, public leaderboards and provider-reported numbers. Independent Artificial Analysis has separately evaluated Argon on its own benchmark suite and gives it an Intelligence Index score of 53.

What is Gemini 4 Argon’s Artificial Analysis score?

Artificial Analysis currently gives Gemini 4 Argon High an Intelligence Index score of 53.

Which Argon benchmark is most impressive?

That depends on use case, but the 84.2% GraphWalks score at 256K-to-1M context, 65.4% Vals Finance Agent score and 77.9% DeepSWE score are among the most practically interesting results.

Bottom Line

Gemini 4 Argon’s benchmark results show a model with unusually broad strength, especially across enterprise knowledge work, long-context reasoning, finance, legal workflows and long-horizon coding.

Google’s most notable published results include:

68.9% — Vals Index

65.4% — Vals Finance Agent v2

19.6% — Harvey’s Legal Agent Benchmark

77.9% — DeepSWE v1.1

84.2% — GraphWalks at 256K–1M

91.7% — LVBench

68% — CWE-bench v1

But Argon also loses important tests.

GPT-6 Astra leads on FrontierSWE v2, Terminal-Bench Science and OSWorld-2.0’s offline subset, while Claude Opus 5.5 leads Terminal-Bench 4.0 and PostTrainBench in Google’s comparison.

Independent Artificial Analysis gives Argon a frontier-level Intelligence Index score of 53, which adds credibility to the broader claim that Google has returned to the top tier of general-purpose AI models.

The most responsible conclusion is not that Argon is universally the best model.

It is that:

Gemini 4 Argon appears to be one of the strongest current models for enterprise knowledge work, long context and several coding workloads, while other frontier models remain stronger on specific execution-heavy tasks.

Benchmarks are a useful filter.

Your own workload should make the final decision.