Cerebras reveals CS-4 with double performance on same chip

Cerebras CS-4 doubles throughput to 4,400 tokens per second without a new chip design.

By Central
The CS-4 achieves double performance by increasing clock speed and adding a third WSE-3 wafer per rack.
Highlights
  • Cerebras doubled the CS-4's performance without designing a new chip by increasing clock speed and improving thermal management.
  • The CS-4 delivers up to 4,400 tokens per second per user, 30 times faster than comparable Nvidia-based systems.
  • The CS-4 is already in production use by OpenAI for its Codex Spark model.

Cerebras has introduced the CS-4 AI accelerator, a system its CEO Andrew Feldman calls the fastest in the industry, and the company has achieved this leap without designing a new chip. The CS-4 is a rack-scale product — a complete server cabinet that integrates compute, power, and cooling into a single data center unit. It still runs on the 5nm WSE-3 chip, the same wafer-scale engine that powered the previous CS-3, but Cerebras has doubled performance by increasing clock speed through more aggressive power delivery and improved thermal management. A single rack now houses three wafers instead of two, and the system delivers up to 4,400 tokens per second per user — a figure Cerebras claims is 30 times faster than comparable setups running on Nvidia GPUs. Memory capacity per wafer remains unchanged at 44 GB. The CS-4 also introduces a modular “Backpack” design for faster assembly and supports disaggregated inference through partnerships with AMD and AWS Trainium. Analysts at SemiAnalysis suggest the networking gains are relatively small, but more technical details are expected at the upcoming Hot Chips conference. The hardware is already in production use by OpenAI for its Codex Spark model, among other customers.

How Cerebras Doubled CS-4 Performance Without a New Chip

The CS-4’s performance gains are a masterclass in system-level engineering rather than a silicon breakthrough. The WSE-3 chip, fabricated on a 5nm process, remains the same massive wafer-scale processor that Cerebras launched in 2024. But the company has found headroom by raising the clock frequency, which requires delivering more electrical power to the chip and extracting the resulting heat far more efficiently. Cerebras does not disclose exact clock speeds or power figures, but the implications are clear: a wafer-scale chip that already consumed enormous power now draws even more, and the cooling system in the CS-4 has been redesigned to handle that load.

Increasing the number of wafers per rack from two to three is another lever. With three WSE-3 chips working in parallel, the CS-4 can process more inference requests simultaneously. This is not a trivial packaging challenge — each wafer is a single, massive die that spans the entire wafer, and integrating three of them into a single rack with shared power and cooling infrastructure required a complete rethinking of the system’s physical layout. The result is a machine that can sustain higher throughput on large language models, a critical advantage for AI inference workloads that demand low latency and high token generation rates.

The token throughput metric — 4,400 tokens per second per user — is particularly striking. To put that in context, many current GPU-based inference systems struggle to deliver a few hundred tokens per second for large models. Cerebras’s 30x claim against Nvidia-based setups is aggressive but plausible if the comparison targets older GPU configurations or models that are not optimized for inference. The CS-4’s architecture is designed from the ground up for dense, compute-heavy inference, whereas GPUs are general-purpose parallel processors that must handle a wider range of workloads.

What is the Cerebras CS-4 and Why Does Rack-Scale Matter?

The CS-4 is a rack-scale AI accelerator, which means it is not a card or a module that plugs into existing servers. Instead, it is a complete, self-contained cabinet that includes the compute wafer, power supplies, cooling systems, and networking — all pre-integrated and tested at the factory. Customers deploy it as a single logical unit, simplifying installation and reducing the complexity of building a large AI cluster from dozens or hundreds of discrete GPU servers.

Rack-scale integration is a defining characteristic of Cerebras’s approach. Most AI hardware vendors sell chips or boards that customers must assemble into servers, network together, and manage with software. Cerebras flips that model: the wafer-scale chip is so large — the size of a full wafer — that it cannot be packaged conventionally. Instead, the company builds the entire system around the chip, treating the rack as the unit of compute. This allows Cerebras to optimize power delivery, cooling, and inter-chip communication in ways that are impossible with discrete components.

For AI inference, rack-scale systems offer a compelling trade-off. They reduce the number of interconnects and the latency penalties that arise when models must be split across many GPUs. The CS-4’s three wafers are tightly coupled, so a large model can be loaded across all three with minimal communication overhead. This is especially valuable for models that are too large to fit on a single GPU but are still manageable on a few wafer-scale processors.

What is the token throughput of Cerebras CS-4?

The Cerebras CS-4 delivers up to 4,400 tokens per second per user. This is the rate at which the system can generate text or other sequential outputs for a single inference request. Cerebras states that this is up to 30 times faster than comparable setups running on Nvidia GPUs, though exact comparisons depend on the model, batch size, and hardware configuration. The throughput is achieved by raising the clock speed of the WSE-3 chip, adding a third wafer per rack, and improving power and cooling to sustain higher performance.

Memory Capacity and the WSE-3 Chip: Consistency Amidst Performance Gains

While the CS-4 doubles performance, memory capacity per wafer remains fixed at 44 GB. This is a notable constraint. The WSE-3 chip integrates its memory on-wafer, meaning there is no separate DRAM pool. The 44 GB is the total SRAM available on the wafer, and it is designed to be extremely fast — it is the memory that the compute units access directly, without the latency overhead of off-chip memory. For inference, this is both a strength and a limitation.

The strength is that the CS-4 does not need to fetch model weights from external memory during inference, which eliminates the memory bandwidth bottleneck that plagues GPU-based systems. The limitation is that any model that requires more than 44 GB of weights cannot fit entirely on a single wafer. Cerebras addresses this by allowing models to be spread across multiple wafers, and the CS-4’s three-wafer configuration provides 132 GB of on-wafer memory. That is sufficient for many large language models, but it still falls short of the memory available in systems with dozens of GPUs that can pool hundreds of gigabytes of HBM.

The WSE-3 chip itself is a marvel of chip design. It contains 4 trillion transistors and 900,000 compute cores, all arranged on a single wafer that is 8.5 inches in diameter. The chip is manufactured by TSMC on its 5nm process, and it is the third generation of Cerebras’s wafer-scale architecture. The decision to keep the same chip for the CS-4 while improving system-level performance suggests that Cerebras sees room for further gains from the WSE-3 without needing a full node shrink or a new architecture. That is a pragmatic strategy, especially given the immense cost and risk of designing a new wafer-scale processor.

The ‘Backpack’ Design and Disaggregated Inference: Cerebras’s New Architecture

Cerebras has introduced a modular “Backpack” design for the CS-4, which speeds up assembly and maintenance. The term “Backpack” refers to a removable chassis that houses the power supplies, cooling loops, and networking components. This module attaches to the main wafer enclosure, allowing technicians to swap out faulty components or upgrade subsystems without dismantling the entire rack. It is a subtle but important engineering improvement: data center operators value systems that can be serviced quickly, and the Backpack design reduces downtime.

More strategically, Cerebras is embracing disaggregated inference through partnerships with AMD and AWS Trainium. Disaggregated inference means that the CS-4 does not need to handle every stage of the inference pipeline alone. Instead, it can work alongside other accelerator types, such as AMD Instinct GPUs or AWS’s custom Trainium chips, for tasks like pre-processing, post-processing, or running smaller models that do not benefit from wafer-scale compute. This is a recognition that AI inference is rarely a single workload; real-world deployments often involve a mix of model sizes and latency requirements. By integrating with other hardware, Cerebras makes its system more flexible and easier to adopt in existing data center environments.

The partnerships also signal a shift in Cerebras’s strategy. Previously, the company marketed its systems as drop-in replacements for GPU clusters. Now, it is acknowledging that heterogeneous compute architectures are the norm, and that the CS-4 can excel as a specialized accelerator for the most demanding parts of the inference pipeline while leaving other tasks to more conventional hardware.

SemiAnalysis and the Networking Question: Small Gains, Big Picture?

Analysts at SemiAnalysis, a respected semiconductor research firm, have noted that the networking improvements in the CS-4 appear to be relatively small. Networking is critical for inference because models must be distributed across multiple wafers, and inter-wafer communication can become a bottleneck. If the CS-4’s networking gains are indeed modest, then the system’s overall performance improvement comes primarily from the higher clock speed and the additional wafer, not from a fundamental advancement in how the wafers communicate.

This is not necessarily a problem. The CS-4 is designed for inference, where the communication pattern is different from training. Inference workloads are often latency-sensitive and require sending a single input through the model, then reading the output. The three wafers can be configured to work in model parallelism — each wafer holds a slice of the model — and the communication required is relatively simple compared to the all-to-all exchanges needed for training. Nonetheless, SemiAnalysis’s observation suggests that Cerebras has not solved the scaling challenge of linking multiple wafer-scale processors. As models grow larger, the company may need to invest in faster interconnects or a different topology.

More details on the CS-4’s networking architecture will likely be revealed at the Hot Chips conference, a key venue for semiconductor and systems companies to present technical details. Attendees can expect to see block diagrams, performance benchmarks, and comparisons with competing systems. The networking specifics will be closely watched because they determine how well the CS-4 can scale to handle models that require more than three wafers.

Cerebras and OpenAI: Real-World Deployment with Codex Spark

The CS-4 is already in production use by OpenAI for its Codex Spark model, an AI coding assistant. This is a significant endorsement. OpenAI is arguably the most demanding customer in the AI industry, and its choice to use Cerebras hardware for inference speaks to the system’s performance and reliability. Codex Spark is a fast, code-generation model that must respond to user queries with low latency; the CS-4’s high token throughput directly benefits the user experience by reducing the time between a prompt and a response.

The partnership also highlights Cerebras’s growing traction in the inference market. Most of the attention in AI hardware has centered on training, but inference is where the majority of compute cycles will eventually be spent. As models are deployed into production, companies care about cost per token and latency, not just raw training throughput. The CS-4’s ability to deliver 4,400 tokens per second per user could translate into lower operational costs for OpenAI, especially if the system can handle multiple concurrent users without degrading performance.

Other customers for Cerebras hardware include research institutions, government labs, and enterprises that need to run large models locally for security or latency reasons. The CS-4 is priced at a premium, but for workloads that demand the fastest possible inference, it may be more cost-effective than building a cluster of hundreds of GPUs.

What to Expect at Hot Chips 2024: More Details on CS-4

The Hot Chips conference, which takes place in August 2024, will be the venue for Cerebras to present the CS-4’s technical architecture in detail. Engineers from the company are expected to disclose the exact clock speed increases, the power consumption of the system, the thermal design, and the specific networking improvements. The conference audience consists of chip architects and system designers, so the presentation will be highly technical.

One area of interest is how Cerebras achieved the clock speed boost. Increasing frequency on a wafer-scale chip is not straightforward because the chip spans the entire wafer, and variations in process, voltage, and temperature across the die can cause timing issues. The company likely had to implement guard bands and sophisticated clock distribution. The cooling solution, too, will be a subject of scrutiny: a three-wafer rack generates enormous heat, and the CS-4 must dissipate it efficiently to maintain stable operation.

Another topic is the “Backpack” design. While Cerebras has described it as a modular maintenance feature, it may also enable future upgrades or variations — for example, a CS-4 configuration that uses different power supplies or networking modules for different data center environments. Hot Chips presentations often include photographs of the system internals, which will give the industry a clearer picture of the CS-4’s physical design.

Finally, the conference will likely feature benchmark comparisons between the CS-4 and competing systems. Cerebras has already claimed 30x performance over Nvidia setups, but independent benchmarks will be needed to validate that figure. If the CS-4 performs well under standard MLPerf inference benchmarks, it could shift the competitive landscape significantly.

The CS-4 is a bold step for Cerebras. By doubling performance on the same chip, the company has demonstrated that system-level innovation can yield substantial gains without waiting for a new silicon generation. At the same time, the fixed memory capacity and modest networking improvements suggest that the architecture has limits. For now, the CS-4 is the fastest inference system in the industry, and it will be a compelling option for organizations that need to serve large language models with minimal latency. Whether that advantage translates into broader market share depends on how well Cerebras can scale the system to even larger models and how quickly competitors respond. The answers will begin to emerge at Hot Chips.

Share This Article