NVIDIA Groq 3 LPX Enters Full Production, Revolutionizing Agentic AI Infrastructure
NVIDIA has announced that its Groq 3 LPX inference accelerator is now in full production, marking a significant step in supporting the next generation of agentic AI systems.
The Groq 3 LPX, designed to complement NVIDIA's Vera Rubin NVL72 AI platform, delivers industry-leading token generation speeds, crucial for applications requiring real-time responsiveness.
In benchmark testing using Gemma 4 31B, an open-source agentic AI model, the Groq 3 LPX produced 3,400 output tokens per second for long-context use cases, a fourfold improvement over competing platforms.
NVIDIA's Vera Rubin NVL72 platform integrates CPUs, GPUs, and NVLink networking into a unified system, offering a scalable solution for agentic AI workloads. By enabling lower token costs and higher efficiency, the platform is designed to meet the needs of hyperscale AI factories, enterprise deployments, and cloud providers.
CoreWeave and Nebius are early adopters of the Vera Rubin NVL72 system and Groq 3 LPX accelerators. SpaceXAI has also announced it will deploy NVIDIA Vera CPUs as part of its next-generation AI architecture, extending the company's capabilities from Earth-based data centers to orbital satellites.
NVIDIA emphasizes 'extreme codesign' as a cornerstone of its approach, co-developing every layer, compute, networking, and inference acceleration, to function as a cohesive unit. This full-stack optimization is particularly relevant as AI inference emerges as a competitive differentiator in large-scale AI deployments.