As AI models grow larger and more complex, especially with trillion-parameter workloads and AI agents becoming mainstream, the demands on AI infrastructure are evolving. Performance now depends not only on raw compute power but also on how compute, memory, storage, networking, and software are integrated into a unified system. NVIDIA’s recent expansion of its NVLink Fusion platform with NVHBM (NVIDIA’s custom high-bandwidth memory) addresses these challenges by enhancing memory performance and efficiency for AI processors, or XPUs.

Traditional high-bandwidth memory (HBM) designs place the memory controller on the XPU chip itself. This approach consumes valuable silicon area that could otherwise be used for compute cores, limiting the overall processing power. NVHBM changes this by embedding NVIDIA’s custom memory controller directly into the 3D HBM stack rather than on the XPU die. This architectural shift results in several key benefits:

  1. Up to 30% higher memory bandwidth, enabling faster data access for AI workloads.
  2. Approximately 15% reduction in power consumption for the memory subsystem, improving energy efficiency.
  3. Frees up to 25% more silicon area on the XPU chip, allowing designers to add more compute resources or other features.

By establishing a standard NVHBM implementation that multiple memory providers can offer, NVIDIA reduces the complexity and engineering effort required to integrate and qualify memory from different suppliers. This standardization helps customers using NVLink Fusion to accelerate the development and deployment of custom AI chips.

A notable collaboration is with Amazon’s Annapurna Labs, which will incorporate NVHBM technology into its next-generation Trainium AI chips. This partnership extends the NVLink Fusion ecosystem, enabling Amazon’s chips and NVIDIA GPUs to work together seamlessly within a common rack-scale architecture. This integration supports more efficient and scalable AI infrastructure designs for hyperscalers and AI innovators.

NVLink Fusion itself is a platform that allows partners to connect custom XPUs and CPUs with NVIDIA’s rack-scale systems. It provides access to NVIDIA’s NVLink chiplets, switches, and MGX systems, along with a broad ecosystem of CPU partners and system manufacturers. This approach lets companies focus on innovating their AI processors while relying on a proven technology stack for networking, system scale-up, and software support.

One practical takeaway is that NVHBM’s architectural innovation—moving the memory controller into the memory stack—demonstrates how rethinking traditional hardware designs can yield significant performance and efficiency gains. For teams building AI infrastructure, this means that memory technology and system integration are just as critical as raw compute power in meeting the demands of next-generation AI workloads.

In summary, NVIDIA’s NVHBM technology within the NVLink Fusion platform represents a strategic advancement in AI infrastructure. By improving memory bandwidth, reducing power consumption, and freeing up chip area, it supports the scaling of AI models and workloads. Collaborations with partners like Amazon further validate this approach and help bring these innovations into practical AI systems.

Read the official announcement (opens in a new tab)

Sources