As trillion-parameter large language models and agentic AI move into the mainstream, the performance bottleneck in AI infrastructure is shifting away from raw compute alone and toward system-level integration across compute, memory, networking, and cooling. NVIDIA has now introduced NVHBM, a new custom high-bandwidth memory technology, and brought it into the NVLink Fusion ecosystem.
Annapurna Labs, Amazon Web Services' chip design arm, is the first partner to adopt the technology. The two sides also agreed on a long-term arrangement covering as many as 2 million top-tier GPUs, pointing to a new relationship between cloud giants and dominant chip vendors that mixes competition with interdependence.
NVHBM shifts how memory and compute space are allocated
In a conventional HBM architecture, the memory controller is typically built into the compute die of the XPU, whether that is a GPU or a custom ASIC. That consumes valuable advanced-process wafer area and limits room for additional compute units.
NVHBM changes that layout. Under NVIDIA's design, the custom memory controller is moved down into and integrated with the base die, the bottom logic die in a 3D HBM stack.
Compared with the standard JEDEC HBM4E specification, ABMedia said NVHBM offers several hardware advantages:
- Up to 30% higher memory bandwidth
- About 15% lower HBM power consumption
- Up to 25% more chip area freed on the XPU compute die, which can be used for additional compute cores or cache memory
NVIDIA is also pushing NVHBM toward a standardized specification across memory suppliers including SK hynix, Samsung, and Micron, a move aimed at lowering the engineering burden for cloud providers buying and validating custom chips.
NVIDIA is trying to pull custom-chip rivals into its own ecosystem
On the surface, cloud service providers have been building in-house ASICs to reduce dependence on NVIDIA and cut capital spending. Rather than answering with a closed system, NVIDIA is using NVLink Fusion and NVHBM to open a path for those rivals to connect into its platform.
ABMedia described three strategic goals behind that move. First, NVIDIA is trying to move beyond being a chip supplier and become a data center architecture setter. Even if customers use internally developed chips, they would still need NVIDIA NVLink chiplets, switches, and MGX rack systems if they build on NVLink Fusion.
Second, the company is using the partnership to counter the UALink open interconnect alliance formed by Broadcom, AMD, and other technology companies. By working with Amazon, NVIDIA is trying to reinforce NVLink's position in data center interconnects.
Third, a unified communication protocol would allow customers to deploy NVIDIA GPUs and in-house ASICs in the same rack, cutting integration barriers between chip architectures and reducing the odds of a full infrastructure shift away from NVIDIA.
AWS puts Trainium4 at the front of the rollout
As the first NVHBM partner, Annapurna Labs said AWS will support both NVLink Fusion and NVHBM starting with its next-generation in-house Trainium4 chip.
At the same time, Amazon confirmed that it plans to buy up to 2 million next-generation high-end GPUs from NVIDIA in 2027 and 2028. The lineup named in the report includes Blackwell Ultra, Rubin, and Rubin Ultra.
ABMedia described the AWS approach as a barbell strategy with two tracks. At the high end, NVIDIA Blackwell Ultra and Rubin series GPUs are aimed at external enterprise customers seeking top-tier performance, as well as frontier model training and complex inference. On the cost-efficiency side, NVLink-enabled Trainium4 is positioned for selected internal workloads and inference tasks where cost matters more.
According to the report, adopting NVHBM lets AWS move Trainium4 forward with lower development risk and higher memory efficiency. It also allows the in-house chip to fit more smoothly into existing server racks and data center networks, reducing friction tied to long-term operations and hardware transitions.
AI infrastructure competition is moving to the system level
The partnership offers a fresh reading of how AI competition is evolving. ABMedia said the choice between building in-house chips and buying NVIDIA hardware is no longer a zero-sum tradeoff, and is instead turning into a mix of architectural integration and functional specialization.
By tying memory architecture through NVHBM to high-speed interconnects through NVLink, NVIDIA is extending its moat from the chip level to the rack-level ecosystem. In ABMedia's framing, the next stage of AI data center competition is moving beyond which chip delivers the most compute and toward who controls the architecture of the AI factory.

