Claude Sonnet 5 beats Opus 4.8 on knowledge work benchmark

Anthropic's mid-tier Claude Sonnet 5 outperforms the larger Opus 4.8 on a knowledge work benchmark, signaling a leap in agentic AI.

By Central
Claude Sonnet 5 achieves 1,618 points on GDPval-AA v2, edging out Opus 4.8's 1,615 in real-world knowledge tasks.
Highlights
  • Claude Sonnet 5 beats Opus 4.8 on the GDPval-AA v2 knowledge work benchmark, scoring 1,618 to 1,615.
  • Anthropic offers introductory API pricing at $2 per million input tokens until August 2026, then rising to $3.
  • Sonnet 5 includes enhanced safety features like default cyber safeguards and reduced vulnerability to prompt injection.

Anthropic has released Claude Sonnet 5, a mid-tier model that closes the performance gap with the company’s flagship Opus 4.8 and, on at least one knowledge work benchmark, surpasses it. Available now across all Claude plans and the API, Sonnet 5 represents a significant leap in agentic capability for Anthropic’s most balanced model family.

How Claude Sonnet 5 Compares to Opus 4.8 and Sonnet 4.6

Anthropic’s published benchmarks show Sonnet 5 beating its predecessor, Sonnet 4.6, in every tested category while narrowing the gap with the pricier Opus 4.8. On agentic coding via SWE-bench Pro, Sonnet 5 scores 63.2 percent, up from 58.1 percent for Sonnet 4.6. Opus 4.8 sits at 69.2 percent. Terminal-Bench 2.1 shows Sonnet 5 at 80.4 percent versus 67.0 percent for Sonnet 4.6. On multidisciplinary reasoning measured by Humanity’s Last Exam, Sonnet 5 reaches 57.4 percent with tools, nearly matching Opus 4.8 at 57.9 percent. Computer use capability, tested on OSWorld-Verified, posts 81.2 percent compared to 78.5 percent for its predecessor.

The standout result comes from the knowledge work benchmark GDPval-AA v2, which tests AI on real-world knowledge tasks. Sonnet 5 actually beats the larger Opus 4.8 here, scoring 1,618 points to Opus’s 1,615. Anthropic reports that feedback from early-access partners confirmed this trend, with the model acting far more agentically than previous versions, particularly in how it handles search tasks.

Safety and Cybersecurity Context for the Sonnet 5 Launch

Anthropic’s release arrives amid heightened scrutiny of its model safety practices. The US government is currently blocking the company’s two most capable models, Mythos 5 and Fable 5, over cybersecurity concerns. Anthropic appears eager to address similar worries with Sonnet 5. The company states the model was not trained on cybersecurity tasks, and in tests for risky capabilities like writing software exploits, it scores far below both Opus 4.8 and Mythos 5.

Sonnet 5 does score slightly higher than its predecessor on these tasks, however. Anthropic has therefore enabled cyber safeguards by default, which flag and block risky cyber usage in real time. These protections are on par with those already in place for Claude Opus 4.7 and 4.8 but are dialed back compared to Fable 5’s guardrails, which users complained about almost immediately. On the broader safety front, the model shows improvements over Sonnet 4.6 in turning down malicious requests, fending off prompt injection attacks, and reducing hallucinations and sycophantic behavior.

Pricing and Availability Through August 2026

Claude Sonnet 5 is live now as the default model for Free and Pro users, and Max, Team, and Enterprise subscribers can access it as well. Developers can integrate it through Claude Code and the Claude Platform. On the API, it is available as “claude-sonnet-5” with a one-million-token context window and a training cutoff of January 2026.

Until August 31, 2026, Anthropic is charging an introductory rate of $2 per million input tokens and $10 per million output tokens. After that, prices increase to $3 and $15, matching what previous Sonnet models cost. A practical caveat: because Sonnet 5 operates more agentically—building plans, grabbing tools like browsers and terminals, and working autonomously—it is likely to consume more tokens per task than its predecessors, potentially raising effective costs even at the same per-token rate.

What This Means for Users and Developers Now

Sonnet 5 delivers a meaningful upgrade for anyone using Claude for knowledge work, coding, or agentic tasks. Free and Pro users can try it immediately as the default experience, while developers on the API can evaluate the introductory pricing before the rate change in September 2026. The key consideration is whether the model’s increased autonomy justifies higher per-task token consumption, but for users who need mid-tier performance approaching flagship levels, Sonnet 5 is available and functional starting today.

Share This Article