NVIDIA Boosts AI Performance with Ray Placement Group Upgrade
NVIDIA has announced an upgrade to its GB300 NVL72 AI platform, enhancing its ability to optimize multi-node GPU scheduling. The update introduces Ray's NVLink Domain-Aware Placement Groups, which can boost AI performance by up to 1.13x.
The GB300 NVL72 is a rack-scale AI platform that integrates 72 Blackwell Ultra GPUs and 36 Grace CPUs per rack. It offers 1.1 exaFLOPS of dense FP4 compute and uses NVLink 5 for all-to-all bandwidth of up to 1,800 GB/s per GPU.
The introduction of NVLink Domain-Aware Placement Groups addresses a limitation in the previous Ray placement groups, which only handled node-level scheduling. This new feature ensures that tightly coupled tasks stay within the same NVLink domain, reducing slower inter-node communication and maximizing the benefits of the high-bandwidth NVLink fabric.
NVIDIA's GEAR research lab tested this feature on large-scale Vision Language Action (VLA) pretraining workloads across a GB300 cluster. The results showed 1.13x faster iterations per second compared to traditional placement methods, demonstrating the potential for significant performance gains in AI applications.