xAI has released Grok 4.5, a large language model trained on tens of thousands of Nvidia GB300 GPUs and designed to compete in coding, agentic tasks, and knowledge work. The model arrives with a pricing structure that undercuts every major rival by a wide margin, raising questions about how much performance matters when the cost gap is this large.
How Grok 4.5 Stacks Up Against GPT-5.5, Fable 5, and Opus 4.8
Benchmark results show a mixed but credible performance picture. On Terminal Bench 2.1, which evaluates complex command-line tasks, Grok 4.5 scores 83.3 percent, nearly matching GPT-5.5 at 83.4 percent and trailing Anthropic’s Fable 5 at 84.3 percent by just one point. That is a strong showing for a first major release in this tier.
The gaps widen on software engineering benchmarks. On DeepSWE 1.1, which measures the ability to resolve real GitHub issues, Grok 4.5 reaches 53 percent, well behind GPT-5.5 at 67 percent and Fable 5 at 70 percent. On SWE Bench Pro, a curated set of harder software engineering problems, the model scores 64.7 percent, beating Opus 4.8 in some configurations but falling short of Fable 5 at 80.4 percent.
These results suggest Grok 4.5 is competitive in terminal-based and agentic workflows but still trails the leaders on complex, multi-step software engineering tasks. The performance is close enough, however, to make the pricing story the real headline.
Grok 4.5 Pricing: A Fraction of the Competition
Grok 4.5 costs $2 per million input tokens and $6 per million output tokens. By comparison, Opus 4.8 charges $5 input and $25 output per million tokens. Fable 5 is at $10 input and $50 output per million tokens. GPT-5.5 and GPT-5.6 sit at $5 input and $30 output per million tokens. The gap is not small — Grok 4.5 is between 60 percent and 80 percent cheaper on output than every model in its performance tier.
xAI also claims that Grok 4.5 uses 4.2 times fewer tokens than Opus 4.8 on SWE Bench Pro tasks, delivering results at 80 tokens per second. Lower per-token pricing combined with fewer tokens per task makes Grok 4.5 by far the cheapest option in this performance tier, assuming the efficiency gains hold up in practice. The strategy mirrors what Chinese vendors like Zhipu and DeepSeek have been doing: get close enough on performance, then win on price.
Training Infrastructure and Data Strategy
xAI says it relied on heavy data filtering, deduplication, and domain-specific selection during training to maintain high data quality. The reinforcement learning stage covered hundreds of thousands of tasks, mostly drawn from software engineering, with automated scoring to evaluate outputs. The training infrastructure was built for asynchronous learning, allowing agentic runs to stretch over many hours while training continued in parallel. These technical choices suggest xAI is optimizing for efficiency and cost control as much as for raw benchmark scores.
Availability, Integrations, and Regional Limits
Grok 4.5 is available now through Grok Build, Cursor, and the xAI console. Plugins are live for Microsoft Word, PowerPoint, and Excel. The model is not yet available in the European Union, with xAI targeting a mid-July launch. xAI trained Grok 4.5 alongside the code editor Cursor, which SpaceX acquired in mid-June for $60 billion in stock, a deal that signals deep integration between the model and the development environment.
What This Means for Developers and Teams
For developers and teams evaluating AI models for coding and agentic tasks, Grok 4.5 presents a compelling cost argument. The model is competitive on terminal-based benchmarks and offers substantial savings on per-token pricing and total token consumption. Teams that can accept slightly lower performance on complex software engineering tasks — or that are willing to test whether the efficiency gains hold in their specific workflows — should try Grok 4.5 through the xAI console or Cursor integration. The model is available right now, and the pricing alone makes it worth a serious look.