Skip to main content
Product Updates

10× Token Capacity: Cerebras Launches WSE-3 Turbo Chip and CS-4 System

Compared with WSE-3, WSE-3 Turbo doubles FP16 sparse AI compute, memory bandwidth, on-chip interconnect bandwidth, and I/O bandwidth across the board, delivering an almost linear performance increase.

10× Token Capacity: Cerebras Launches WSE-3 Turbo Chip and CS-4 System

Cerebras, a manufacturer of wafer-scale AI inference accelerators, announced on the 18th local time the launch of its next-generation WSE-3 Turbo chip and the CS-4 system based on it. Compared with the previous-generation CS-3, CS-4 delivers 10× the token capacity and 2× the speed.

10× Token Capacity: Cerebras Launches WSE-3 Turbo Chip and CS-4 System

Like the previous WSE-3, WSE-3 Turbo integrates 900,000 cores and 44GB of on-chip SRAM cache, but its FP16 sparse AI compute, memory bandwidth, on-chip interconnect bandwidth, and I/O bandwidth have all doubled. The CS-4 system can integrate three WSE-3 Turbo chips, further increasing compute density.

Cerebras says CS-4 achieves a token delivery rate of 4465 Token/s on the OpenAI GPT-OSS model, 30 times that of GPU solutions, while improving by 93% over CS-3. For ultra-high-parameter models, CS-4’s 2μs low-latency wafer-to-wafer interconnect enables more than 1000 Token/s of token output at the 10T scale.

10× Token Capacity: Cerebras Launches WSE-3 Turbo Chip and CS-4 System

CS-4 natively supports splitting inference workloads and can serve as a decoding unit paired with heterogeneous prefill hardware, leveraging the differentiated architectural strengths of both.

This computing hardware is based on Cerebras’s all-new Nexus rack-scale platform. Its computing, power, and interconnect subsystems have been redesigned as independent systems, reducing the total number of components by 50% and substantially increasing manufacturing automation. CS-4 nearly eliminates board-level power delivery losses and provides WSE-3 Turbo with twice the power supply.