NVIDIA expands NVLink Fusion with NVHBM custom memory technology

NVIDIA is introducing NVHBM, a custom high-bandwidth memory technology designed to increase performance and efficiency for third-party AI accelerators.

Image: NVIDIA Developer

NVIDIA has announced the expansion of its NVLink Fusion platform with the introduction of NVHBM, a custom high-bandwidth memory technology aimed at hyperscalers and AI-native companies. This development is designed to support the creation of semi-custom AI accelerators, or XPUs, by providing a standardized, validated memory architecture that integrates directly into the NVIDIA AI infrastructure ecosystem.

The core innovation of NVHBM lies in its architectural shift regarding the memory controller. Traditional HBM designs place the controller on the XPU die, which consumes significant silicon area. By moving the memory controller into the 3D HBM stack itself, NVHBM frees up to 25% more area on the compute die, allowing designers to allocate more space for matrix engines, vector units, and other workload-specific capabilities.

Beyond area savings, the technology offers measurable performance improvements over standard HBM4e. According to NVIDIA, NVHBM delivers up to 30% higher memory bandwidth and 15% lower power consumption. These gains are intended to keep compute engines consistently fed with data, which is critical for memory-bound AI workloads such as large-model inference and training.

NVIDIA is positioning NVHBM as a standardized offering available through multiple memory providers. By providing a validated base-die technology, the company aims to reduce the engineering complexity and supply chain bottlenecks typically associated with qualifying custom memory for new accelerator programs. This approach allows partners to leverage NVIDIA’s existing scale-up and scale-out technology stack, including the MGX rack-scale architecture.

Amazon’s Annapurna Labs has been named as the first partner to collaborate on NVHBM technology. This partnership builds upon previous efforts to integrate NVLink Fusion into AWS infrastructure. Annapurna Labs plans to utilize this architecture in its upcoming Trainium4 chips, enabling Amazon’s custom silicon to operate within a unified rack-scale environment alongside NVIDIA GPUs.

Sources

  1. NVIDIA DeveloperNVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure
  2. NVIDIA BlogNVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory