Vera Rubin Delivers Blistering Performance Gains, But Upgrades Come with Hefty Price Tag
Nvidia's next-generation Vera Rubin platform has delivered impressive results in AI performance benchmarks. The latest MLPerf Inference v6.1 tests showed that Vera Rubin NVL72 achieved up to 3.7x the throughput of Nvidia's GB300 NVL72 on the Qwen3-VL benchmark, and 2.5x on DeepSeek-R1.
Nvidia described these results as a preview submission ahead of broader deployment, which means that this performance jump is specific to certain workloads. The company also noted that four GB300 NVL72 racks containing 288 GPUs achieved 99% scaling efficiency, and software improvements alone produced performance gains of up to 1.6x over the previous MLPerf round.
The benchmark results highlight an unusually fast replacement cycle in AI infrastructure. Cloud operators are repeatedly deciding whether existing clusters remain competitive against newer systems capable of producing substantially more tokens per rack and per megawatt.