Cerebras Systems has introduced the CS-4, a rack-scale accelerator that packs three of the company's new WSE-3 Turbo wafers into a single system rated at 750 PFLOPS of AI compute. Announced on August 18 from the company's Sunnyvale headquarters, the machine is aimed squarely at inference workloads where response speed, rather than raw training throughput, decides whether a product feels usable. First units ship this quarter.
The headline claim is interactivity at scale. On GPT-OSS-120B, Cerebras reports more than 4,400 output tokens per second for a single user, which it frames as up to 30 times the per-user speed of leading GPU deployments and roughly double its own CS-3. Supporting that are 129.6 petabytes per second of memory bandwidth and 7.2 terabits per second of system I/O, up from 1.2 terabits on the previous generation.
A rack rebuilt around the wafer
CS-4 is the first product on what Cerebras calls the Nexus Platform Architecture, which splits the rack into three independently serviceable layers: compute, power and I/O. Compute arrives as a Wafer-Scale Backpack, a sealed assembly folding the wafer, power conversion, direct liquid cooling, high-speed I/O and control electronics into one 3D package with half the component count of the previous design. Because the PowerRack can be installed and facility-qualified before any compute shows up, Cerebras says installation drops from days to hours.
Two engineering choices carry most of the performance story. Power conversion now sits 0.5 millimeters from the processor instead of the roughly 50 millimeters typical of a GPU board, nearly eliminating board-level loss and letting the company push twice as much power into each wafer at higher clock frequencies. A new programmable I/O subsystem with RoCE v2 and what Cerebras calls Direct Wafer Links then lets wafers talk to one another inside and across racks without a switch, cutting wafer-to-wafer latency from five microseconds to as low as two.
That latency number is the one that matters for very large models. Cerebras says CS-4 sustains more than 1,000 tokens per second on models above 10 trillion parameters, and that the fabric is designed to address models beyond 50 trillion parameters, a scale no deployed system has reached but one the company is clearly reserving runway for. It also claims up to 10 times more throughput per watt than CS-3, an efficiency argument aimed at data-center operators rather than benchmark watchers.
Reading the numbers carefully
Every figure here originates with Cerebras or with benchmarking it commissioned, and the company's own footnote concedes that observed gains vary by workload, configuration and model tested. Single-user token rates also flatter architectures that keep an entire model resident in fast on-wafer memory, while the metric hyperscalers actually purchase on is aggregate throughput per dollar across thousands of concurrent sessions. Cerebras has not published head-to-head figures on that second axis.
The commercial context still argues for taking the launch seriously. Cerebras hardware powers OpenAI's Ultrafast mode for GPT-5.6 Sol, the coding platform Lovable has taken dedicated capacity for latency-sensitive workloads, and the company has guided to roughly $880 million to $890 million in core revenue for 2026 following its public listing. Nvidia remains the default purchase for nearly every buyer, and a wafer-scale rack is a heavier commitment than adding another GPU node.
What CS-4 tests is whether speed has become a product feature rather than an infrastructure detail. As agents chain dozens of model calls into a single user request, latency compounds in a way that a faster rack can flatten. Cerebras is betting a meaningful slice of the market will pay for the fastest tokens rather than the most familiar ones.






