Nvidia Rebuilds Memory Around Amazon's Chip

Nvidia's blog announced NVHBM on August 26, 2026, describing a new architecture that moves Nvidia's custom memory controller into the high-bandwidth memory base die itself, rather than leaving it on the compute chip. Nvidia's post said the design delivers up to 30 percent greater memory bandwidth and 15 percent lower HBM power consumption than commodity HBM4e, while freeing up to 25 percent more area on the compute die for other logic. The technology is being made available to partners in Nvidia's NVLink Fusion program, the licensing scheme that lets outside chipmakers connect their own silicon to Nvidia's interconnect fabric. Tom's Hardware independently corroborated the same three figures after reviewing Nvidia's technical briefing, describing the shift as Nvidia exporting its memory-controller design past the boundary of its own chips for the first time.

Amazon's Annapurna Labs is named as the first NVHBM partner. Nvidia's blog post quoted Nafea Bshara, vice president at Annapurna Labs, saying NVHBM represents a new architectural approach to advancing high-bandwidth memory performance and efficiency. The collaboration extends a commitment AWS made in December 2025, when it agreed to build NVLink Fusion into its custom silicon roadmap; ServeTheHome's coverage of that earlier announcement described Trainium4 as designed from the outset to integrate with NVLink and Nvidia's MGX rack architecture. Annapurna Labs says it will support NVLink Fusion starting with Trainium4, letting AWS-designed accelerators and Nvidia GPUs share a common rack-scale architecture.

MetricNVHBM vs. commodity HBM4e
Memory bandwidthUp to 30 percent higher
HBM power drawUp to 15 percent lower
Freed compute-die areaUp to 25 percent more

The Chip Amazon Sold As the Way Around Nvidia

Trainium was not pitched to customers as a complement to Nvidia GPUs. TechCrunch's tour of Amazon's Trainium lab, published in March 2026, described Trainium explicitly as an alternative to Nvidia's backlogged, hard-to-acquire GPUs, with AWS staff saying Trainium3 UltraServers cost up to 50 percent less to run for comparable performance than classic cloud servers. AWS also cut switching costs on purpose: moving a workload from Nvidia hardware to Trainium was described as basically a one-line PyTorch change and a recompile. One AWS director told the outlet that the chip was breaking records on price per power, part of what the piece called an attempt to chip away at Nvidia's market dominance wherever possible.

That positioning is now harder to square with Trainium4's own architecture. The chip Amazon built to reduce Nvidia dependency will ship using Nvidia's interconnect standard and, through NVHBM, memory built to Nvidia's own controller specification. Nvidia does not need to sell AWS a single GPU for this arrangement to matter: it earns a licensing and ecosystem role inside the chip that was designed to replace its products. The independence Trainium offered was about avoiding Nvidia's silicon and its supply queue. NVHBM shows how much of Nvidia's underlying architecture Trainium4 will still carry regardless.

Why a Clean Break From Nvidia May Not Exist

The obvious question beyond this one announcement is whether any credible de-Nvidia compute strategy is left to buy into. Google's TPUs and Microsoft's Maia chips reduce direct GPU purchases the same way Trainium does, but none of them operate in a vacuum: as Nvidia opens NVLink Fusion and now NVHBM to outside silicon, it is positioning its interconnect and memory architecture as an industry default rather than a Nvidia-only feature. A custom chip can still cut a buyer's GPU bill. It is a separate question whether that chip's performance now assumes Nvidia's memory controller design and Nvidia's scale-up fabric are available and licensed on Nvidia's terms.

This is not the same as Trainium becoming an Nvidia product; Amazon still designs, owns and prices the chip. But independence from Nvidia was never a single switch to flip, and it looks less like one after this deal than it did before it. What Trainium4 shows is that 'custom silicon' and 'Nvidia-free' have quietly become different claims, and a buyer choosing between AWS's own chip and an Nvidia GPU should not assume the two options are architecturally separate any more.

What EU and UK Compute Buyers Should Track

For a European operator choosing AWS custom silicon specifically to hedge Nvidia pricing power or GPU allocation risk, the practical question is no longer just which chip logo sits on the rack. It is which company's interconnect and memory architecture the deal actually runs on, and whether that dependency shows up in the contract, the roadmap, or neither. Procurement and infrastructure teams evaluating multi-cloud or multi-vendor compute strategies should ask providers directly which generation of interconnect and memory underlies the capacity being quoted, rather than inferring it from the chip's brand name.

Three risks are worth tracking on a standing basis rather than revisiting only at contract renewal: single point of failure, since one company's interconnect and memory design increasingly sits underneath chips marketed as alternatives to that same company; pricing power, since Nvidia can influence HBM economics even on hardware it does not sell; and supply concentration, since NVHBM still depends on the same small group of memory manufacturers already supplying Nvidia's own accelerators. None of these three risks are unfamiliar to procurement teams. They now apply, less obviously, to the very chip an EU or UK buyer may have picked specifically to avoid them.