Anthropic Launches Claude Opus 5, Cheaper Model for Coding and Agents

Anthropic releases Claude Opus 5, a cost-effective AI model that delivers near flagship performance at half the price, targeting enterprise coding and agent tasks.

By Central
Claude Opus 5 offers near flagship performance at half the cost, targeting enterprise tasks.
Highlights
  • Claude Opus 5 delivers nearly the full capability of Anthropic's flagship Fable 5 at half the cost.
  • On the Frontier-Bench v0.1 coding benchmark, Opus 5 scored 43.3 percent, more than double its predecessor.
  • Anthropic now stratifies its lineup into Fable 5, Opus 5, Sonnet 5, and Haiku 4.5 for different tasks.

The artificial intelligence industry has spent years obsessing over what the most powerful models can achieve in controlled tests. Anthropic’s launch of Claude Opus 5 signals a decisive shift toward a more pragmatic, economically driven era of competition, where the measure of a model is not just its peak intelligence but the price of deploying that intelligence at scale. Released across all of Anthropic’s platforms, Opus 5 is positioned as the company’s new “daily driver,” delivering nearly the full capability of its flagship Claude Fable 5 at precisely half the cost, a move that redefines the battleground for enterprise AI adoption.

Anthropic Stratifies Its Lineup with the Opus 5 “Daily Driver”

Anthropic is not claiming that Claude Opus 5 is its smartest model; that distinction still belongs to Claude Fable 5. Instead, the company is making a deliberate strategic argument about the nature of economically valuable AI work. Most enterprise tasks do not require the absolute frontier of intelligence. They require a model that is exceptionally capable, reliable, and crucially, cheap enough to run on every task, every day. Opus 5 is designed to fill this exact niche.

An Anthropic spokesperson described the new lineup as a stratified toolkit for different classes of work. Fable 5 is reserved for the most ambitious, days-long autonomous projects. Sonnet 5 is for work run at high scale, where speed and cost per call are the deciding factors. Haiku 4.5 is optimized for subagents and instant answers. Opus 5 sits squarely in the middle, described as “the model you hand complex work to and review when it’s done.” This framing represents a maturity in how AI companies are thinking about product-market fit, moving away from a one-size-fits-all approach toward specialized tools for specific economic functions.

The core distinction Anthropic is drawing is between “bounded tasks” and “long-horizon autonomy.” Bounded tasks have a specific outcome and a clear finish line, such as debugging a function, generating a report from structured data, or building a feature from a specification. Long-horizon autonomy involves extended, multi-step projects that require a model to stay coherent over hours or days. Opus 5, Anthropic argues, is the best tool for the jobs that benchmarks can see and measure. Fable 5 is what you reach for when the job outruns the benchmark entirely.

How Claude Opus 5 Benchmarks Against Fable 5 and Rival Models

The benchmark results Anthropic released paint a picture of a model that is not just cost-effective but genuinely competitive at the frontier on key practical tasks. On Frontier-Bench v0.1, an agentic terminal coding benchmark, Opus 5 scored 43.3 percent, more than double the 18.7 percent achieved by its predecessor Opus 4.8 and significantly ahead of Fable 5’s 33.7 percent. On ARC-AGI 3, an evaluation of novel problem-solving, Anthropic reports Opus 5 scored three times higher than the next best model. On the OSWorld 2.0 computer-use benchmark, the company says Opus 5 surpasses Fable 5’s best result at just over a third of the cost.

These numbers come with unusually candid caveats from Anthropic. The company acknowledges that Opus 5 remains behind a competing model, Mythos 5, on cybersecurity tasks and biology research. An OpenAI-family model still holds the lead on one agentic coding benchmark. This honesty is notable in an industry prone to superlatives and suggests a confidence in the overall value proposition that goes beyond benchmark victories.

What distinguishes Claude Opus 5 from the more expensive Claude Fable 5? Opus 5 is optimized for discrete, economically significant tasks where efficiency and cost-per-task are paramount. It excels at work that has a clear start and finish, such as coding a specific function or analyzing a defined dataset. Fable 5, on the other hand, is designed for extended, multi-day autonomous projects requiring sustained coherence across many connected steps. The choice between them depends entirely on whether the workload is a bounded task or a long-horizon job.

The more revealing insight came when the Anthropic spokesperson was asked directly where Opus 5 still falls short of Fable 5. Their answer cut to the heart of what modern benchmarks do and do not capture. The evals where Opus 5 wins are bounded tasks with a specific outcome. What those evals do not measure is duration. Fable 5 is for the longest, most autonomous jobs, where the model has to stay coherent across many connected steps over hours or days with dense source material. The spokesperson advised customers to run both models on a representative workload, one bounded task and one long-horizon job, to see which fits their needs. This framing of bounded tasks versus long-horizon autonomy may become the defining axis of model differentiation in the coming year, as benchmarks saturate and the hardest remaining problems involve sustained, multi-day agentic work rather than discrete puzzles.

The Real Battleground: Token Efficiency and Enterprise AI Economics

Threaded through the entire launch announcement is a single, dominant theme that Anthropic clearly wants enterprise buyers to absorb: Opus 5 does not just score well, it scores well per dollar. The model ships with an adjustable “effort” setting that allows customers to trade peak intelligence for speed and token savings. The company’s charts emphasize performance at a given cost rather than peak performance alone, a subtle but profound shift in how AI capability is being marketed.

Early customer testimonials reinforce this point with concrete efficiency data. Niko Grupen, head of applied research at the legal AI company Harvey, reported that Opus 5 achieved similar performance to Opus 4.8’s maximum-reasoning mode while generating 26 percent fewer tokens on average. Richard Pham of Fundamental Research Lab said that on hard financial-modeling tasks, the model averaged nine percentage points higher accuracy while using roughly one-third fewer turns and tool calls and 60 percent less time. Wade Foster, CEO of Zapier, said Opus 5 topped his company’s AutomationBench leaderboard without spending more tokens than prior Claude models, running a full churn-prevention workflow from start to finish. Scott Wu, CEO of Cognition, the company behind the Devin coding agent, said that on FrontierCode 1.1, Claude Opus 5 approaches Fable-level performance at half the cost, with particular strength in debugging and root-cause analysis.

This efficiency emphasis reflects a commercial reality. Enterprise AI spending is no longer experimental. Inference costs, the price of actually running these models at scale, have become a board-level line item. Anthropic’s business skews heavily toward API and enterprise usage. Analysis from early 2026 indicated that Claude held roughly 40 percent of the enterprise large language model market by usage, and Claude Code alone had reached about $1 billion in annualized revenue. For a company whose customers pay by the token, a model that does more with fewer tokens is not a nice-to-have. It is the product itself.

Building Agents That Check Their Work: The Self-Verification Advantage

Beyond the numbers, Anthropic is selling a behavioral story about a model that verifies its own work and iterates until it succeeds. The company offered several examples from internal testing that read like small parables of machine stubbornness and self-correction.

In one Frontier-Bench task, the model was asked to reconstruct a machine part as a 3D CAD model from a drawing it was intentionally given no way to view. Rather than fail, Opus 5 wrote its own computer vision pipeline to extract the geometry from raw pixels, solving the task repeatedly while no competing model succeeded in five attempts. In another case, given a real bug in a popular open-source package manager, the model found the root cause and fixed an edge case that the community’s own patch had missed. A competing model patched only the symptom and declared victory. An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Finding no live feed to validate against, the model built its own test harness to check its parsing code.

Customers described similar behavior in production. Cristian Rivera, a staff software engineer at Stripe, said he gave the model a chief-of-staff role over his development environments for a weekend. The model built its own monitor, drove each server, and pulled him in only for the judgment calls. This is the capability enterprises actually care about. The gap between a model that produces plausible output and one that verifies its output is the gap between a demo and a deployable system. Most of the hidden cost of enterprise AI today is human review, engineers checking the machine’s work. A model that reliably checks its own work compresses that cost, which is precisely why customers keep citing fewer turns, fewer passes, and less time rather than higher raw benchmark scores.

Safety Architecture in the Opus 5 Era: Capability Gaps and Classifier Fallbacks

The launch also showcases Anthropic’s increasingly intricate approach to safety, one that now involves deliberately not teaching its models certain skills. Anthropic says its automated behavioral audit found Opus 5 to be its most aligned model to date, scoring 2.3 on overall misaligned behavior, lower than Opus 4.8, Sonnet 5, or Fable 5, with the lowest rates of deceptive behavior and the least susceptibility to being tricked into misuse.

On the capability side, Anthropic intentionally avoided training Opus 5 on cyber tasks, as it did with Opus 4.8. The model improved on them anyway, a side effect of general capability gains, and now nearly matches Mythos 5 at finding software vulnerabilities. However, it remains far behind at exploiting them. On Anthropic’s OSS-Fuzz evaluation, Opus 5 identified vulnerabilities at a 79.4 percent rate, close to Mythos 5’s 80 percent, but succeeded at developing exploits in only 4 challenges versus Mythos 5’s 13. This asymmetry, strong at defense-relevant discovery and weak at offense-relevant exploitation, appears to be by design.

The safeguards follow the same logic. Anthropic expects Opus 5’s cyber classifiers to intervene about 85 percent less often than Fable 5’s. When a classifier does trigger, requests in Claude.ai, Claude Code, and Claude Cowork fall back to Opus 4.8 by default. This raises an obvious question: if a request is too risky for one model, why is it acceptable for another? The model it falls back to has lower capability levels, making the risk of harmful use lower as well. The logic is defensible, revealing how AI safety works in practice: risk is not a property of the question alone, but of the question multiplied by the capability of the system answering it. On biology, the calculus runs the other way. Opus 5 is now Anthropic’s most capable generally available model for scientific research, scoring 10.2 percentage points higher than Opus 4.8 on the company’s internal chemistry benchmark, though the company acknowledges that Mythos 5 remains stronger for long-horizon, open-ended work like autonomous drug design campaigns.

The Business and Regulatory Context of the Opus 5 Launch

The launch lands at a moment of extraordinary commercial momentum for Anthropic. The company was valued at roughly $380 billion in its latest funding round. Annualized revenue climbed from about $1 billion at the end of 2024 to a projected $9 billion by the end of 2025, with internal targets reportedly reaching $20 to $26 billion for 2026. These targets are underwritten by enormous infrastructure commitments, including a reported $30 billion Azure compute deal alongside arrangements with Google Cloud and Nvidia. This level of spending only pencils out if enterprises keep expanding their usage of Anthropic’s models.

That is the context in which Opus 5’s pricing strategy makes sense. Holding the price at Opus 4.8 levels while roughly doubling performance on key agentic benchmarks is effectively a steep price cut per unit of capability, designed to widen the funnel of workloads that are economical to automate. Every task that was marginal at Opus 4.8’s cost-per-success becomes viable at Opus 5’s, and every viable task is recurring token revenue.

The regulatory backdrop has grown more complex. A U.S. judge gave final approval to a $1.5 billion copyright settlement with book authors, closing a chapter of litigation over the company’s early training data. Separately, the U.S. government moved to block foreign access to Anthropic’s most advanced models, a reminder that frontier AI is now entangled with export policy in ways that shape which customers can buy what. Also shipping with Opus 5 is a Fast mode running at roughly 2.5 times default speed at twice the base price, automatic fallback routing on the API, and mid-conversation tool changes that no longer invalidate the prompt cache, a small feature that agent developers may appreciate more than any benchmark. Consistent with prior Opus models, Opus 5 carries no data retention requirements for general access.

Two questions will determine whether the bet pays off. The first is whether Opus 5’s efficiency claims survive contact with production workloads at scale. The second is whether enterprises embrace a world where safety classifiers, not users, sometimes decide which model answers a question. The deeper message of this launch is that the AI industry’s center of gravity has moved. For years, the labs competed on what their best model could do on its best day. With Opus 5, Anthropic is competing on something less glamorous and far more lucrative: what a very good model can do every day, for half the price. In a market where the frontier keeps moving, the real fortune may lie just behind it, in the realm of reliably excellent, economically sustainable intelligence.

Share This Article