Google Gemini 4 Argon matches rivals on key benchmarks but trails competitor lead

Google's Gemini 4 Argon re-enters the frontier AI race, matching rivals on benchmarks but still trailing Anthropic's lead.

By Central
Gemini 4 Argon ties with rivals on the AI index at 53 points but trails Claude Opus 5.5 at 58.
Highlights
  • Gemini 4 Argon scores 53 on the Artificial Analysis Intelligence Index, tying with GPT-6 Astra and Claude Fable 5.1.
  • The model achieves first place on AutomationBench-AA with 77.5 percent, ahead of Claude Sonnet 5.5.
  • Google's Gemini 4 Argon lags behind Anthropic's Claude Opus 5.5 which leads the index at 58 points.

Google has re-entered the frontier AI race with Gemini 4 Argon, its first top-tier model in over seven months, and the numbers show a company that has closed the gap with OpenAI and Anthropic without quite seizing the lead. The model matches rivals on several key benchmarks, undercuts them on price, and even surpasses them in human preference rankings for text generation. Yet questions about efficiency, real-world software integration, and whether Google can maintain this momentum linger beneath the otherwise favorable data.

Seven Months of Silence and a Skipped Generation

The road to Gemini 4 Argon was anything but smooth. Google had already announced—and then abandoned—Gemini 3.5, skipping an entire frontier generation amid what insiders described as a difficult and drawn-out development period. That gap left Gemini 3.1 Pro as the company’s flagship for more than half a year, a stretch during which OpenAI released GPT-6 Astra and Anthropic delivered both Claude Opus 5.5 and Sonnet 5.5. The pressure on Google’s DeepMind division was mounting, and the decision to leap directly to Argon reflects a strategic pivot: skip the incremental update and aim for a model that can credibly compete at the highest tier.

Argon now puts Google back among the top three AI labs by most independent measures, though Anthropic likely retains the overall lead.

Argon now puts Google back among the top three AI labs by most independent measures, though Anthropic likely retains the overall lead. The question for developers, enterprises, and consumers is not whether Argon is good—it clearly is—but where it excels, where it still lags, and whether the price advantage holds up under real workloads.

What Independent Benchmarks Reveal About Argon’s True Position

Artificial Analysis, an independent evaluator, provides one of the earliest third-party assessments. At its highest available reasoning setting, “High,” Gemini 4 Argon scores 53 points on the Artificial Analysis Intelligence Index. That score ties it with OpenAI’s GPT-6 Astra (max) and Anthropic’s Claude Fable 5.1, and places it one point ahead of GPT-6.1 Sol (max). The improvement over Google’s previous frontier model, Gemini 3.1 Pro Preview, is a striking 23-point jump.

Yet the top of the index remains occupied by Anthropic. Claude Opus 5.5 leads at 58 points, followed by Claude Sonnet 5.5 at 56. Argon sits in a tie for third, which is a credible position for a model that arrived months after its competitors, but it is not a victory lap. The model’s performance on specialized agentic benchmarks tells a more nuanced story. On AutomationBench-AA, an agentic task suite from Artificial Analysis, Argon takes first place with 77.5 percent, six points ahead of Claude Sonnet 5.5 (max). That represents a significant gain for Google, whose previous models had known weaknesses in agentic workflows. On Terminal Bench 4, Argon scores 57 percent—a 53-point improvement over Gemini 3.1 Pro Preview—but still trails Claude Sonnet 5.5 (64 percent), Claude Opus 5.5 (60 percent), and GPT-6 Astra (59 percent).

Hallucination Rates and Accuracy: A Tradeoff Worth Understanding

One of Argon’s most distinctive characteristics is its unusually low hallucination rate. On the AA-Omniscience benchmark, which tests factual knowledge and a model’s willingness to admit ignorance, Argon hallucinates only 15 percent of the time. That compares favorably with GPT-6 Astra at 51 percent and GPT-6.1 Sol at 54 percent. In practice, this means Argon is far more likely to respond with “I don’t know” than to fabricate an answer. For enterprises in regulated industries such as finance, law, and healthcare, this trait alone could justify adoption.

The tradeoff is accuracy. Argon reaches only 50 percent on the same benchmark, five points below Gemini 3.1 Pro Preview and 13 points below GPT-6 Astra (max, 63 percent). When Argon does attempt an answer, it gets it right less often than its competitors do. The overall score on the Omniscience benchmark reflects this balance: Argon lands at 42 points, roughly even with GPT-6 Astra (43) and GPT-6.1 Sol (42). The model is more honest but less confident, and that is a deliberate design choice rather than a flaw. Google has clearly prioritized reducing harmful fabrications over maximizing correct answers, a tradeoff that aligns with the company’s stated focus on safety and trustworthiness.

Human Preference Rankings: Argon Takes the Text Crown

On Arena.ai, where human raters compare model outputs head-to-head, Gemini 4 Argon (High) achieves 1,525 points in the Text Arena, placing it first overall and 20 points ahead of Claude Opus 4.6 (High) in second. This is a notable result because human preference rankings measure something that benchmarks cannot: the subjective quality of the output as experienced by real users. Google’s previous model, Gemini 3.8 Flash (High), had been languishing in eleventh place.

According to Arena, Argon leads across multiple categories including coding, hard prompts, instruction following, longer queries, and creative writing. It also ranks first across all evaluated professional fields and for queries in English, Chinese, and Russian. This broad strength in human evaluation suggests that Google has invested heavily in alignment and output quality, not just raw benchmark performance.

The web development results are more modest. In Code Arena: WebDev, Argon scores 1,679 points and lands in eighth place. That is a 96-point improvement over Gemini 3.8 Flash (High) and a jump from 29th position, but it does not crack the top tier dominated by specialized coding models.

Pricing Strategy: Cheap Tokens, Expensive Habits

Google has set an introductory price of $2 per million input tokens and $10 per million output tokens for Gemini 4 Argon. Cached input tokens cost 95 percent less, working out to approximately 10 cents per million. Regular pricing will double those rates to $4 input and $20 output. Compared to GPT-6 Astra at $10 input and $50 output, or Claude Opus 5.5 at $4 input and $20 output, Argon looks aggressive on raw token pricing.

However, token price is only one side of the cost equation. Argon consumes an average of 62,000 output tokens per task in the Artificial Analysis Intelligence Index, while GPT-6 Astra uses only 27,000. At the current promotional rate, one Intelligence Index task costs $1.99 with Argon, about 60 percent of GPT-6 Astra’s $3.26. Once the discount ends, that cost rises to $3.98, approximately 20 percent above GPT-6 Astra. The model’s lower token rates do not compensate for its higher token consumption. The price advantage is real but temporary, and it depends heavily on workload patterns. For tasks that benefit from long, chain-of-thought reasoning, Argon’s token consumption may be justified. For straightforward queries, the efficiency gap is a genuine cost concern.

Price per million tokens Gemini 4 Argon (promotional) Gemini 4 Argon (regular) GPT-6 Astra Claude Fable 5.1 Claude Opus 5.5
Input Token $2 $4 $10 $10 $4
Output Token $10 $20 $50 $50 $20
Cache Read $0.10 $0.20 $1 $0.25 $0.20
Cache Write N/A N/A $12.50 $12.50/up to $20 $5/up to $8

One Million Token Output: A New Capability for Long Reasoning

Google has raised the output limit from 64,000 tokens to one million tokens, calling it an industry first. The idea is straightforward: if a model can generate hundreds of thousands of tokens in a single trajectory, it can reason through complex problems more thoroughly and solve them in one pass rather than requiring multiple calls. To support this capability, Google has added a “Long Decode Continuation” feature to the Gemini API, which pauses long responses and resumes them through follow-up requests so that the reasoning process does not hit a timeout.

The input context window remains at one million tokens. Argon accepts text, images, video, and audio as input but outputs only text. For enterprises dealing with long-form document analysis, legal briefs, or scientific research papers, the expanded output window could be a genuine differentiator. No competing frontier model currently offers a comparable output limit, and this alone may drive adoption among users who need sustained, coherent reasoning over extended generation tasks.

Internal Benchmarks and the Vals Index: Google’s Own Numbers

Google’s internal benchmark results paint a significantly more favorable picture than some independent evaluations. The company’s published data shows Argon leading across most tested categories, sometimes by wide margins. The model also tops the Vals Index at 68.9 percent, making it the first Gemini model to achieve first place on that ranking. Argon finishes in the top five on 20 of 22 tested benchmarks, with particular strength in finance, law, coding, and security.

Some caution is warranted. The same Vals Index shows that Anthropic’s mid-tier model, Claude Sonnet 5.5, also outranks its own top-tier Opus 5.5, suggesting that the index may not perfectly capture real-world utility. Internal benchmarks from any AI lab should be read as directional rather than definitive. Independent evaluations from sources like Artificial Analysis and Arena.ai provide a more reliable basis for comparison.

Rollout and Access: The Phased Approach

Gemini 4 Argon is not yet available to the general public. Google has adopted a phased rollout strategy that begins with a group of “trusted cyber defenders” through the Fairwind program. These early testers, along with Google’s internal teams, will receive the model without cyber guardrails, allowing them to stress-test its capabilities in real-world security scenarios. Google describes this as a necessary step given the capabilities of frontier models and cites participation in the US government’s voluntary AI testing program as part of its justification.

Feedback from these early testers will feed into the model’s safety mechanisms before wider release. Google plans to open Argon to developers, businesses, and consumers after that phase, starting with paying API customers and Google AI Ultra subscribers. The company has not provided a specific date, stating only that access will come “as soon as possible.” For developers and enterprises evaluating whether to build on top of Argon, this timeline introduces uncertainty. Committing to a model that may change its safety profile or availability window carries inherent risk.

The Software Gap: Where Google Still Trails

Independent evaluations of the model itself tell only part of the story. As has been repeatedly demonstrated in the AI market, a model’s real-world performance depends not just on its architecture and training but on the software ecosystem wrapped around it. Google’s Gemini app currently lags behind Anthropic’s Claude Cowork and OpenAI’s ChatGPT Work in terms of user experience, tooling, and integration.

Argon may be a strong model, but users interact with software, not raw APIs. Google has not announced any major updates to its consumer or enterprise applications alongside the Argon release. If the company does not invest equally in the user-facing layer, the model’s benchmark advantages may never translate into market advantage. The same dynamic has hurt Google in previous AI product cycles: strong research capabilities that fail to produce compelling end-user experiences.

For now, Gemini 4 Argon is best understood as a credible return to form rather than a definitive victory. It matches or beats rivals on several important metrics, offers a compelling price point at launch, and introduces genuinely novel capabilities in long-form generation. But it also consumes more tokens per task than its competitors, trails Anthropic on overall intelligence index scores, and faces a software ecosystem that is not yet competitive with the leaders. The next six months, as Argon reaches more users and the initial promotional pricing period ends, will determine whether Google has truly rejoined the frontier or simply posted strong numbers in a race it is not yet equipped to win.

Questions answered
  • How does Gemini 4 Argon perform on independent AI benchmarks?Gemini 4 Argon scores 53 points on the Artificial Analysis Intelligence Index, tying with GPT-6 Astra and Claude Fable 5.1, while Claude Opus 5.5 leads at 58 points.
  • What are Gemini 4 Argon's strengths in agentic tasks?The model takes first place on AutomationBench-AA with 77.5 percent, outperforming Claude Sonnet 5.5 by six points.
  • What are the limitations of Gemini 4 Argon?Gemini 4 Argon trails Claude Opus 5.5 on the overall intelligence index and consumes more tokens per task compared to competitors.
Share This Article