Nvidia Claims Fourfold Speed Boost with Groq 3 LPX Chip Amid Scalability Concerns
Nvidia has announced that its Groq 3 LPX chip has entered full production. The company acquired the Groq license for about $20 billion in late December and claims it is four times faster than the Cerebras chip, but some experts warn that this comparison may be stacked in Nvidia's favor.
The Groq 3 LPX is an interactive AI inference accelerator designed to deliver ultrafast token generation for agentic AI systems. According to Nvidia, its speed allows agents to iterate more often within a reasonable timeframe, making coding tasks take 'minutes instead of hours'. A benchmark from Artificial Analysis shows the chip reaching 3,400 tokens per second on the Gemma 4 31B model with a 100,000-token context window.
The comparison to Cerebras is based on performance held steady between 10,000 and 100,000 tokens of input length. However, experts point out that Groq relies on an SRAM-heavy dataflow architecture with limited memory per LPU (Local Processing Unit), requiring models to be split across several accelerators over Ethernet.
Nvidia plans to offer the chip through its Token Factory and is among the early users of the technology. The chip's scalability with larger mixture-of-experts models remains an open question, with some models potentially needing hundreds or even thousands of LPUs to run efficiently.