Anime, manga, and games, with a take · A Yukimedia publication

← all stories otherannouncement 1 sources · 1h ago ·

Cerebras CS-4 Rack System Runs AI Inference Up to 30 Times Faster Than GPUs

The CS-4's claimed 30x inference speed advantage over GPUs, combined with a 10x improvement in throughput per watt over the CS-3, directly targets data center profitability by delivering more tokens within a limited power budget.

Reporting from 1 source: GIGAZINE.

Cerebras CS-4 Rack System Runs AI Inference Up to 30 Times Faster Than GPUs

Cerebras unveiled the CS-4 rack-scale solution, claiming it performs AI inference up to 30 times faster than simple GPU systems. The system features three Wafer Scale Engine 3 Turbo processors per system, with each wafer achieving up to twice the speed of the previous generation. Initial shipments begin this quarter.

Cerebras has introduced the CS-4, a rack-scale system built around three Wafer Scale Engine 3 Turbo processors per unit. The company says each wafer runs up to twice as fast as the prior generation, and the new power, cooling, and I/O architecture draws more performance out of each wafer.

The system claims inference speeds up to 30 times faster than simple GPU setups, and can process over 1,000 tokens per second on models exceeding 10 trillion parameters. Cerebras states the CS-4 improves throughput per watt by up to 10 times compared to its predecessor, the CS-3.

By placing the power supply 0.5mm from the processor, roughly 100 times closer than on traditional GPU boards, the design reduces power loss and doubles the power that can be supplied. The company says deployment time drops from days to hours, with initial shipments beginning this quarter.

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources