Cerebras CS-4 Server Smashes Nvidia GPUs in AI Chatbot Queries
Cerebras Systems has unveiled its CS-4 server rack, touting that it can deliver up to 30 times more tokens per second per user than Nvidia GPUs when running large language models.
The new hardware is powered by three WSE-3 Turbo processors, described as the largest AI semiconductors ever built. Each processor packs 4 trillion transistors and enables the system to produce over 4,400 tokens per second per user on the GPT-OSS-120B model.
This represents a significant improvement from its previous CS-3 predecessor, delivering double the speed while achieving up to 10 times higher throughput per watt.