NVIDIA Groq 3 LPX Smashes AI Inference Performance Records
NVIDIA's Groq 3 LPX has shattered performance records in AI inference, reaching an unprecedented 3,431 tokens per second on a 100K context benchmark. This achievement was made possible by NVIDIA's advanced Vera Rubin platform and the system's focus on deterministic execution, low-latency token generation, and fine-grained scheduling.
The Groq 3 LPX, built around NVIDIA's LP30 accelerators and rack-scale architecture, boasts 315 PFLOPS of FP8 inference compute and 128 GB of SRAM. Its ability to handle high-interactivity and long-context AI workloads is particularly crucial for agentic AI tasks, such as coding or reasoning workflows.
NVIDIA reports that the Groq 3 LPX can generate 5,000 tokens in just 1.5 seconds, compared to 50 seconds at 100 TPS - an order-of-magnitude leap in efficiency. The system's performance on a 10K context benchmark was also impressive, delivering 3,382 TPS with minimal variation in latency.