Updated 27 Aug 2026: Adds Cerebras' Aug 25 Hot Chips roadmap: Nexus is explicitly designed to span CS-4, CS-5 and CS-6; CS-5 is targeted for 2027 with up to 10,000 output tok/s on named open models, while CS-6 adds wafer-scale 3D DRAM integration. All forward-looking performance targets remain clearly attributed.
Cerebras CS-4 turns Nexus into a multi-generation rack-scale inference platform
Cerebras’ fourth-generation system is more than a faster wafer: CS-4 debuts the modular Nexus rack architecture, and an August 25 roadmap now makes clear that Nexus is intended to carry CS-5 and CS-6 rather than end with one generation.
CS-4 was announced August 18, 2026 as Cerebras’ fourth-generation rack-scale system.
Each CS-4 contains three WSE-3 Turbo processors; Cerebras says each wafer provides 250 PFLOPS of AI compute, 43.2 PB/s of memory bandwidth and 2.4 Tb/s of off-wafer I/O.
CS-4 is the first implementation of the modular Cerebras Nexus Platform Architecture, separating compute, power and I/O into independently evolving subsystems.
On August 25, Cerebras said Nexus is designed to support multiple generations and was co-designed for the next-generation WSE that will debut in CS-5.
CS-5 is targeted for 2027; Cerebras says it is designed for up to 10,000 output tokens per second per user on selected open models and up to 5,000 on the largest frontier models. These are forward-looking vendor targets.
Cerebras says CS-6 will add wafer-scale 3D integration with stacked DRAM to increase memory capacity while preserving locality.
The Wafer I/O Module supports RoCE v2 RDMA over Ethernet and Direct Wafer Links, with claimed wafer-to-wafer latency as low as two microseconds.
Cerebras explicitly supports disaggregated inference in which other accelerators perform prefill and CS-4 performs decode.
Cerebras says first CS-4 shipments begin in Q3 2026.
What builders should take away
Inference providers evaluating Cerebras can now model Nexus as a multi-generation platform rather than a one-off CS-4 rack, which matters when planning refresh cycles and integration work.
For agentic, coding, voice or interactive reasoning products, benchmark end-to-end latency and completed-task throughput; peak tokens-per-second numbers do not capture prompt prefill, tool calls or network overhead.
Infrastructure teams should test RoCE interoperability, rack power/cooling requirements and operational tooling before assuming Nexus’ modularity translates directly into easier deployment in their environment.
Treat Cerebras’ current 30× GPU-speed claim and its CS-5/CS-6 roadmap figures as vendor claims until independently reproduced on relevant workloads.
Watch provider pricing: the architectural advance matters to most builders only when CS-4 capacity appears through accessible inference services with competitive economics.
What changed
On August 18, 2026, Cerebras introduced CS-4, a rack-scale system built from three WSE-3 Turbo processors and the first implementation of its new Nexus Platform Architecture. Nexus separates compute, power and I/O into modular assemblies. On August 25 at Hot Chips, Cerebras added an explicit multi-generation roadmap: Nexus is designed to support CS-4, CS-5 and CS-6; CS-5 is targeted for 2027 with a next-generation WSE, while CS-6 is planned to combine wafer-scale SRAM and compute with 3D-stacked DRAM. Cerebras says this modular platform gives it a path to roughly double token-generation speed year over year for the next several years. These roadmap numbers are forward-looking vendor targets, not independent benchmarks.
Why it matters
For builders and infrastructure operators, CS-4 changes the Cerebras story from a single unusual accelerator into a more modular rack-scale platform, and the Hot Chips roadmap strengthens that interpretation by showing Nexus is intended to persist across multiple processor generations. The new I/O and disaggregated-inference support make it easier to consider Cerebras alongside existing GPU or ASIC capacity rather than as an all-or-nothing stack. If Cerebras’ deployment, latency and future throughput claims hold in production, the architecture could materially alter the economics and responsiveness of reasoning-heavy, agentic and other latency-sensitive AI services. Procurement decisions still need workload-specific validation because the headline performance figures and future-generation targets remain vendor-produced.
Nexus makes the rack itself the product architecture
CS-4 is the first system built on Cerebras’ Nexus Platform Architecture. Instead of treating compute, power and I/O as one tightly coupled rack design, Nexus packages them as modular subsystems. The compute element is a rear-mounted Wafer-Scale Backpack containing the wafer, power conversion, direct liquid cooling, high-speed I/O and controls. Cerebras says this design has 50% fewer components than its prior system and uses 60% more automated manufacturing, with the goal of reducing installation from days to hours.
Hot Chips turns Nexus into an explicit multi-generation roadmap
On August 25, Cerebras said Nexus was co-designed for its next-generation WSE and previewed both CS-5 and CS-6. CS-5 is targeted for 2027 and is designed, according to Cerebras, to reach up to 10,000 output tokens per second per user on named open models such as Gemma 4 31B and gpt-oss-120b, and up to 5,000 output tokens per second per user on the largest frontier-class models. CS-6 is planned around 3D integration, pairing wafer-scale SRAM and compute with stacked DRAM so more model state can remain close to the wafer. These are forward-looking product targets rather than shipping specifications.
Power and I/O changes are central to the performance story
Cerebras moves power conversion to about 0.5 mm from the processor, which it says cuts board-level loss enough to deliver nearly twice as much power at almost the same voltage. The new Wafer I/O Module doubles off-wafer bandwidth to 2.4 Tb/s per wafer and supports both RoCE v2 RDMA over Ethernet and Direct Wafer Links. Cerebras says the direct links reduce wafer-to-wafer latency to as low as two microseconds, which is intended to preserve interactive decode as models span more wafers and racks.
CS-4 is designed for heterogeneous inference, not only Cerebras-only clusters
A material architectural addition is explicit support for disaggregated inference. Cerebras describes GPU or ASIC systems handling prompt prefill and CS-4 handling low-latency decode, with AMD Helios and AWS Trainium named as complementary platforms. RoCE connectivity gives operators a standards-based integration path, while the programmable I/O layer leaves room for more specialized interconnects.
The speed claims need to be separated from the architecture facts
Cerebras advertises up to 30× faster inference than production GPU systems, up to 10× more throughput per watt than CS-3 and more than 1,000 tokens per second on models exceeding 10 trillion parameters. Its Hot Chips roadmap adds even larger CS-5 targets and an architectural CS-6 preview. Cerebras itself labels the roadmap forward-looking, and its site notes that performance comparisons rely on third-party benchmarking or internal testing. Builders should therefore treat the architecture and announced roadmap as established facts, while treating current and future performance numbers as claims to validate.
Availability starts with infrastructure buyers
Cerebras says first CS-4 shipments begin in the third quarter of 2026. That makes this initially an infrastructure and provider story rather than a direct API change for most application developers. The downstream significance will depend on how quickly cloud and inference providers expose CS-4 capacity, what they charge, and whether lower latency survives real multi-tenant production workloads.
What to watch next
Independent benchmarks of CS-4 across representative models, long contexts, concurrency levels and full request latency.
Which cloud, neocloud and inference providers deploy CS-4, and their pricing and service-level terms.
Production evidence for heterogeneous prefill/decode deployments using AMD, AWS or other accelerators with CS-4.
Whether Nexus’ modular rack design materially shortens manufacturing, deployment and upgrade cycles in customer data centers.
Whether Cerebras hits the 2027 CS-5 schedule and whether the claimed generation-over-generation token-speed gains survive production deployment.
Further technical detail on CS-6's planned 3D-stacked DRAM integration, capacity, thermals and availability.
Still unclear
Most performance and efficiency figures are Cerebras claims based on internal testing, selected third-party benchmarks or extrapolation; independent CS-4 production data is still limited.
CS-5 and CS-6 details are explicitly forward-looking and may change before release.
Cerebras has not published enough public pricing to compare CS-4 total cost of ownership directly with current GPU racks across common workloads.
The practical complexity and economics of disaggregated prefill/decode depend on software, networking and workload characteristics that are not fully described in the launch material.
GPT-5.6 Sol Ultrafast remains in limited preview, but OpenAI’s August 21 standard-tier price cut changes its economics: Sol input is now 20% cheaper and output 33% cheaper through at least November 21. Ultrafast pricing is still undisclosed.
Groq 3 LPX is moving from architecture announcement to manufactured infrastructure. Artificial Analysis measured about 3,400 output tokens/s at both 10K and 100K context on an NVIDIA-hosted private endpoint, but the single-concurrency benchmark does not yet establish public-cloud price, multi-tenant throughput or end-to-end agent speed.
Jalapeño is working first-party silicon rather than a roadmap item, and OpenAI now says AI itself materially accelerated the design process. The distinction still matters: tape-out means the design was finalized for manufacturing; it does not mean fleet-scale production qualification or API deployment is complete.