Nvidia Vera Rubin NVL72 Hits Full Production with ExaFLOPS Racks
Nvidia's next data-center platform, Vera Rubin NVL72, has entered full production. The company confirmed this in its own newsroom statement, saying it is 'ramping into full production' with Taiwan's top server makers and global supply chain leaders manufacturing Vera Rubin-based systems at scale.
Vera Rubin NVL72 pairs a new Rubin GPU with a new Vera CPU, shipping as a rack-scale system rather than a single chip. The platform was first shown publicly at GTC 2025 and used to confirm production in the CES 2026 keynote. Nvidia's newsroom announced it had moved past sampling and qualification straight into full-scale manufacturing.
The NVL72 rack packs 72 Rubin GPUs and 36 Vera CPUs into one liquid-cooled cabinet, tied together by nine NVSwitch 6 blades across 18 compute blades. The fabric delivers 3.6 TB/s of bandwidth to each GPU, twice the per-GPU interconnect bandwidth of the prior-generation rack platform.
Nvidia's own product page reports a single NVL72 rack at roughly 3,600 PFLOPS of NVFP4 inference performance and 2,520 PFLOPS of NVFP4 training performance. This would put one rack well into exaFLOPS territory for the low-precision math formats that dominate large language model inference today.