Google develops Frozen v2 chip for Gemini, with projected efficiency 6x to 10x above TPU

Google develops Frozen v2 chip for Gemini, with projected efficiency 6x to 10x above TPU

N
News Editor
2026-07-21 04:36:32
Google is developing a new server chip code-named Frozen v2 to improve inference efficiency for Gemini, according to a report by The Information. The chip’s main design feature is that parts of the Gemini model architecture would be embedded directly into silicon hardware, reducing both the number of computations needed to answer user prompts and the amount of data that must move inside the chip. Google engineers estimate that, under the same power consumption, the chip could process 6x to 10x as many tokens per watt as the company’s current Tensor Processing Units, or TPUs. The company is aiming for deployment in 2028, though the reported performance figure remains an estimate rather than a measured benchmark. The report also said Google does not plan to manufacture Frozen at TPU scale. Frozen v2 is being treated in part as a technical experiment, will likely remain for internal use, and is not expected to become an external product because its design is tied closely to a specific Gemini architecture. The project is also linked to Google’s internal compute shortage, which has reportedly caused friction inside the company and forced Google Cloud to turn away some customer business.
GoogleGeminiFrozen v2TPUAI chipscomputeAlphabet

Google is developing a new server chip, code-named Frozen v2, to improve Gemini inference efficiency. According to The Information, Google engineers estimate that, at the same power draw, the chip could process 6x to 10x more tokens per watt than the company’s existing Tensor Processing Units, or TPUs. The target for deployment is 2028. After the report emerged, Alphabet shares closed up 1.52% on Monday.

Gemini architecture partially embedded in silicon

The defining feature of the chip, the report said, is that part of the Gemini model architecture would be embedded into the silicon hardware itself. That would cut the number of computations needed when the chip answers user prompts, while also reducing the volume of data that has to move within the chip. The result is higher inference efficiency.

The name Frozen refers to the idea of “freezing” part of the model’s computational logic into silicon. Google engineers’ 6x to 10x projection refers to token throughput per watt relative to TPU under the same power consumption. The report noted that this figure is still an estimate, not measured performance.

Not a TPU replacement and not planned for mass production

The Information said Google is not planning to manufacture Frozen at the same scale as TPU. The chip is not intended to replace TPU, which serves broader use cases, and Google currently sees part of Frozen v2 as a technical prototype.

Its long-term usefulness also depends on whether Google keeps the same underlying architecture for future Gemini models. If Gemini undergoes major architectural changes, the chip could quickly become outdated. Because the design is tied to a specific architecture, Frozen v2 is not expected to be sold as an external product.

Compute shortage is part of the reason behind the project

The project is partly aimed at easing a serious internal compute shortage at Google. The report said the shortfall has created friction inside the company and has even forced Google Cloud to decline some customer business.

Last month, Google agreed to pay SpaceX nearly $1 billion per month to help fill that compute gap.

Google’s AI business is also facing broader pressure

Beyond compute constraints, Google’s AI push is dealing with other challenges. The report said the release timeline for the next Gemini Pro has slipped, several senior researchers have been poached by rivals, and Chinese models are gradually taking a share of the US enterprise market. New models released over the weekend by Moonshot AI and Alibaba have also drawn market attention.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
600

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.