Google Cloud: AI Is Memory-Bound Now, HBM Tops 75% of AI Server BOM

Google Cloud: AI Is Memory-Bound Now, HBM Tops 75% of AI Server BOM

N
News Editor
2026-09-02 02:17:35
At SEMICON Taiwan 2026, Google Cloud senior director Nikhil Cherian said AI workloads are now memory-bound, with high-performance memory exceeding 75% of AI server hardware BOM cost. Google unveiled a two-pronged approach: TPU 8i (288GB HBM, 384MiB on-chip SRAM, zero off-chip latency) and TPU 8t (9,600 chips, 2PB pooled HBM) to split inference and training, plus a training-free lossless quantization algorithm called TurboQuant that compresses KV cache from 32-bit to 3-bit, cutting memory footprint sixfold and accelerating attention computation 8x. The algorithm also integrates older DRAM technology to extend component lifecycle.
SEMICON Taiwan 2026's memory forum kicked off on September 1, and Nikhil Cherian, senior director of supply chain infrastructure at Alphabet's Google Cloud, used the stage to make a blunt point: AI compute has shifted from being compute-bound to memory-bound. High-performance memory now accounts for more than 75% of the bill of materials (BOM) cost on AI servers, he said. To work around the capacity, bandwidth and power bottlenecks, Google is attacking the problem from both sides of the stack. On the hardware side, it is splitting inference and training workloads across different chips. On the software side, it is using lossless quantization to shrink memory footprint.

Two TPUs, two jobs

Google's hardware strategy starts with a new pair of TPUs. The TPU 8i targets low-latency inference, packing 288GB of high-bandwidth memory while tripling on-chip SRAM capacity to 384MiB. By moving the dynamic conversation state and key-value cache onto the chip itself, the TPU 8i avoids off-chip latency entirely. For large-scale training, the TPU 8t strings together 9,600 chips into one massive compute cluster. HBM pooling extends to 2PB, which removes the off-chip data-transfer bottleneck. The design also works with Google's TPU Direct Storage technology.

TurboQuant: squeezing KV cache from 32 bits to 3

On the software front, Google has developed TurboQuant, a training-free, lossless quantization algorithm. It compresses the key-value cache of large models from 32 bits down to 3 bits, cutting memory usage by a factor of six with no loss in accuracy. The algorithm also delivers an 8x speedup in attention computation. TurboQuant additionally pulls in older-generation DRAM integration technology to extend the life cycle of components.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.