Google Delivers Eighth-Gen TPU: Splits Training and Inference into Two Chips, Claims 80% Better Performance Per Dollar

Google Delivers Eighth-Gen TPU: Splits Training and Inference into Two Chips, Claims 80% Better Performance Per Dollar

N
News Editor 01
2026-07-22 12:08:13
At Cloud Next 2026, Google unveiled its eighth-generation TPU, splitting training and inference into dedicated TPU 8t and TPU 8i chips. The inference chip claims 80% better performance per dollar over Ironwood. Anthropic and OpenAI have already adopted the TPU cluster.
Google TPUeighth-generation TPUtraining inference splitAI chipNvidia competition

At Cloud Next 2026, Google introduced its eighth-generation TPU with an unprecedented move: splitting AI training and inference tasks across two dedicated chips. The TPU 8t is designed for training, while the TPU 8i focuses on inference. Google claims the latter delivers 80% better performance per dollar than the previous-generation Ironwood.

Why Split Training and Inference?

AI compute involves two fundamentally different phases. Training requires extreme compute density to learn from massive datasets. Inference handles real-time queries after deployment, demanding low latency and low cost. Previously, a single TPU tried to do both. From the eighth generation onward, they are separate.

The TPU 8t delivers 12.6 petaFLOPS of FP4 compute, 216 GB of high-bandwidth memory, and 6.5 TB/s memory bandwidth. Google says it is 3x faster than the previous generation for training and can scale to over one million TPUs in a single cluster.

The TPU 8i provides 10.1 petaFLOPS of FP4 compute, 288 GB HBM, and 384 MB on-chip memory to reduce data movement latency. Google claims an 80% improvement in inference performance per dollar at low latency targets. Both chips are expected to be publicly available by the end of 2026.

Challenging Nvidia's Moat

Nvidia currently uses a single GPU product line for both training and inference. Its upcoming Vera Rubin offers 35 petaFLOPS FP4, 288 GB HBM4, and 22 TB/s bandwidth — raw compute numbers still far ahead of the TPU 8t's 12.6 petaFLOPS. But the real battleground is cost per inference. The TPU 8i is engineered to lower unit costs, which is exactly what large model providers like Anthropic and OpenAI care about.

Notably, Anthropic announced it is expanding Claude's training and serving to a multi-gigawatt TPU capacity, making it the largest known TPU customer. OpenAI has also started using Google TPU capacity.

Yet Google is not abandoning Nvidia. It announced that its cloud will offer Nvidia Vera Rubin chips by the end of 2026. The two companies are jointly enhancing the Falcon network protocol, an open-source data center networking technology from Google, to make Nvidia systems run more efficiently on Google Cloud.

The Ecosystem Challenge and Cloud Provider Dynamics

The real challenge for Google TPU has never been just specs. Nvidia owns CUDA, a software framework that developers have relied on for two decades. Training scripts, optimization tools, and academic reproducibility all depend on CUDA. Google TPU has its own compiler toolchain, but every workload migration creates friction for engineers.

Amazon and Microsoft face the same structural difficulty: they each develop custom chips to reduce reliance on Nvidia while continuing to buy Nvidia hardware. The logic of hyperscalers is not to kill Nvidia, but to capture more profit on workloads where their own chips are effective. By splitting training and inference, Google sends a clear engineering signal: it is serious about competing on inference costs — not replacing Nvidia, but making sure customers have no reason to keep paying Nvidia's premium in specific scenarios.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.