Vera Rubin NVL72 Crushes GB300 in Benchmark Debut
Nvidia's next-generation AI rack, Vera Rubin NVL72, has made its first public appearance in benchmark data. In the MLPerf Inference v6.1 round, released on September 16, 2026, Vera Rubin showed up to 3.7x the throughput of Nvidia's current-generation GB300 platform.
The results were buried in a 486-result dataset and are considered a preview submission, as the hardware and software stack are not yet shipping configurations. However, this debut is significant because it provides the first independently verified data point on Vera Rubin's real-world inference performance.
Nvidia used a full 72-GPU NVL72 rack for its submission, with individual Rubin accelerators listed under the designation VR200. The company disclosed a larger four-rack configuration, 288 GPUs total, running DeepSeek-R1 offline at 99% scaling efficiency.