NVIDIA Unveils NVLink 6 for Large-Scale AI Deployments
NVIDIA has launched NVLink 6, its latest interconnect technology designed to support large-scale AI deployments. Built within the Rubin platform, NVLink 6 delivers a multi-layer resiliency framework to ensure continuous operation and uptime in massive GPU clusters.
The new technology boasts a bidirectional bandwidth of 3.6 TB/s per GPU, doubling the performance of its predecessor, NVLink 5, and offering over 14 times the bandwidth of PCIe Gen6. A single 72-GPU Rubin NVL72 rack equipped with NVLink 6 can achieve an unprecedented 260 TB/s of all-to-all bandwidth.
NVLink 6's approach to fault tolerance spans the physical, link, and application layers, using lightweight Forward Error Correction (FEC) and Physical Layer Retry (PLR) mechanisms to correct bit-level errors without introducing significant latency. The technology also supports advanced recovery features like NVIDIA's Shadow Engine Recovery, which can restore inference operations in seconds.
NVLink 6 addresses the growing infrastructure demands of AI factories, where maintaining cluster utilization is critical. As AI models grow larger and deployments scale to thousands of GPUs, maintaining continuous uptime is essential for revenue generation in inference-heavy workloads.