Nvidia Executives Confirm CPX Cancellation and Detail 2026 Groq LPU Integration Roadmap

By Central

In the wake of Nvidia’s landmark GTC 2026 keynote, where CEO Jensen Huang unveiled the Vera Rubin GPU architecture and the strategic acquisition of Groq’s LPU technology, the company’s Vice President of Hyperscale and HPC, Ian Buck, provided critical clarifications on the future of its data center roadmap. In an exclusive press Q&A, Buck confirmed the shelving of the anticipated CPX platform and outlined a concrete timeline for the integration of Groq’s Language Processing Unit (LPU) into Nvidia’s ecosystem.

The Strategic Pivot Away from CPX

One of the most significant revelations from the session was the formal confirmation that Nvidia’s CPX (Coherent Processor Accelerator) project has been shelved. This move represents a major strategic pivot for the company’s data center ambitions. Buck explained that the decision was driven by a confluence of architectural evolution and market feedback.

“Our analysis, combined with deep engagement with hyperscale partners, showed that the Vera Rubin architecture’s advancements in native tensor memory and inter-GPU coherence reduced the need for a separate, discrete coherence accelerator like CPX,” Buck stated. “The performance targets we had for CPX are being met and exceeded through the integrated design of Vera Rubin and our next-generation NVLink fabric.”

Market Forces and Technological Convergence

Buck elaborated that the rapid adoption of CXL (Compute Express Link) 3.0 as an industry-standard cache-coherent interconnect further diminished the unique value proposition of a proprietary solution like CPX. The decision to shelve CPX allows Nvidia to reallocate significant engineering resources towards its core GPU and systems roadmap, including the full-stack integration of the newly acquired Groq LPU technology. This shift underscores a broader industry trend where specialized, standalone accelerator cards are being subsumed into more holistic, system-level architectures that offer greater efficiency and programmability.

Groq LPU Integration: A Timeline for Decode Acceleration

The focal point of the discussion quickly turned to Nvidia’s acquisition of Groq and its LPU technology. Buck provided the first concrete details on how and when this technology will reach customers. Contrary to speculation about a lengthy integration period, Buck announced that the first fruits of this acquisition will be available within the current calendar year.

“We are on track to ship our first LPU-powered decode acceleration solutions by the end of 2026,” Buck revealed. “This isn’t about rebadging existing Groq hardware. We are deeply integrating the LPU’s deterministic tensor streaming architecture into our NVIDIA AI Enterprise software stack and our HGX platform. The goal is to offer a seamless, high-efficiency solution for the inference bottleneck that is large language model decoding.”

Architectural Synergy and Customer Benefits

The integration plan highlights a synergistic approach. The Vera Rubin GPU will handle the computationally intensive “prefill” phase of generative AI—processing the initial prompt—while the dedicated LPU will take over the “token generation” or decode phase. This bifurcation leverages the strengths of each architecture: the massive parallel compute of the GPU for the single, large batch operation of prefill, and the ultra-low latency, deterministic execution of the LPU for the sequential, autoregressive task of generating tokens one by one.

Buck emphasized that for customers, this will manifest as a dramatic increase in tokens-per-second for LLM inference at a significantly lower total cost of ownership (TCO). “We’re targeting an order-of-magnitude improvement in inference throughput and latency for sustained conversational AI and code generation workloads,” he said. The solution will be offered both as a software update for compatible existing systems and as a new hardware configuration for future HGX server nodes.

The Q&A session also touched upon the broader implications of managing a portfolio that now includes three distinct processing architectures: the traditional GPU, the Data Processing Unit (DPU) from the Mellanox acquisition, and now the LPU. Buck was adamant that this is not a challenge, but a strategic advantage.

Unified Software: The Common Denominator

“The key is NVIDIA AI Enterprise and our CUDA ecosystem,” Buck asserted. “The developer experience is unified. A developer writes to our APIs for inference, and the software stack intelligently schedules workloads across the available GPU, LPU, and even CPU resources within the node. The complexity of orchestrating these diverse engines is our responsibility, not the customer’s.”

This software-centric approach is designed to prevent fragmentation and ensure that the addition of the LPU simplifies the deployment of large-scale AI rather than complicating it. Buck pointed to the success of integrating DPUs for networking and security offload as a blueprint for the LPU’s integration for inference offload.

Competitive Landscape and Industry Response

When pressed on the competitive landscape, where companies like AMD, Intel, and a host of startups are also pursuing AI accelerator strategies, Buck pointed to Nvidia’s full-stack approach as its defining moat. “Others may offer a chip. We offer a data center-scale computing platform—from the silicon to the systems to the software to the services. The LPU is a powerful new tool in that toolbox, making our platform even more capable and efficient for the generative AI era.”

Industry analysts view this move as a preemptive strike to solidify Nvidia’s dominance in AI inference, a market segment that is poised for explosive growth as trained models move into widespread production. By neutralizing a potential competitive threat in Groq and absorbing its differentiating technology, Nvidia simultaneously removes a rival and strengthens its own product lineup.

The confirmation of the CPX cancellation and the detailed LPU roadmap from Ian Buck provides unprecedented clarity on Nvidia’s strategic direction post-GTC 2026. The company is decisively moving away from auxiliary accelerator projects in favor of deepening the integration and capabilities of its core platforms. The promise of shipping integrated LPU decode acceleration by year’s end signals a new phase in the AI hardware wars, where efficiency and total system performance for production inference become the primary battlegrounds, with Nvidia positioning its full-stack platform as the most comprehensive solution available.

Share This Article