Nvidia Unveils Custom High-Bandwidth Memory to Boost AI Performance
Nvidia has unveiled its custom high-bandwidth memory architecture, NVHBM, which promises to deliver up to 30% more memory bandwidth per stack and 15% lower HBM power consumption compared to standard HBM4E.
The new architecture moves the memory controller off the accelerator die and into the base die of the 3D stack, reducing interface and support area by up to 67% against the JEDEC HBM4E standard. This improvement is expected to address growing performance limitations for frontier-level AI models centered around limited bandwidth.
Nvidia's NVHBM is a building block gated behind its NVLink Fusion program, which connects third-party accelerators to its rack-scale platform. Amazon's Annapurna Labs is the only named partner so far and has not explicitly mentioned NVHBM, but will support NVLink Fusion with Trainium4.
The technology is not new, as Marvell announced a similar idea in December 2024, claiming up to 25% more compute area, 33% greater memory capacity, and a 70% reduction in memory interface power. Nvidia's approach reimagines where memory is managed, rather than reinventing how it is handled.