Nvidia's $20 Billion LPU Bet Pays Off with Impressive Benchmarks
Nvidia's $20 billion gamble on Groq's LPU technology seems to be paying off, with impressive benchmarks showing its LPX rack systems can process 3,400 tokens per second (tok/s) using Google's Gemma 4 31B model.
This is a significant speedup over the nearest alternative platform, Cerebras, which managed only 882 tok/s under the same conditions. The Groq 3-based LPX systems use an SRAM-heavy dataflow architecture, which provides much faster memory bandwidth than traditional GPUs.
However, the LPUs have limited onboard memory, so Nvidia's architecture distributes models across multiple accelerators using Ethernet. Each LPX rack can be equipped with up to 256 LPUs for a total of 128 GB of high-bandwidth SRAM.