Groq 3 LPX Shatters AI Inference Speed Records with 3,431 TPS Benchmark
NVIDIA's Groq 3 LPX has set a new standard in AI inference, achieving an impressive 3,431 tokens per second on a 100K context benchmark.
Benchmarked by Artificial Analysis using the Gemma 4 31B model, this performance highlights the system's ability to handle high-interactivity and long-context AI workloads on NVIDIA's advanced Vera Rubin platform.
The Groq 3 LPX is built around NVIDIA's LP30 accelerators and rack-scale architecture, offering 315 PFLOPS of FP8 inference compute and 128 GB of SRAM. Its focus on deterministic execution, low-latency token generation, and fine-grained scheduling enables it to excel in multiturn agentic sessions and interactive workloads.
The system's performance was particularly crucial for agentic AI tasks, such as coding or reasoning workflows, where models must process large volumes of accumulated context. NVIDIA reports that the Groq 3 LPX can generate 5,000 tokens in just 1.5 seconds, compared to 50 seconds at 100 TPS, an order-of-magnitude leap in efficiency.