Google is developing a new server chip, code-named Frozen v2, to improve Gemini inference efficiency. According to The Information, Google engineers estimate that, at the same power draw, the chip could process 6x to 10x more tokens per watt than the company’s existing Tensor Processing Units, or TPUs. The target for deployment is 2028. After the report emerged, Alphabet shares closed up 1.52% on Monday.
Gemini architecture partially embedded in silicon
The defining feature of the chip, the report said, is that part of the Gemini model architecture would be embedded into the silicon hardware itself. That would cut the number of computations needed when the chip answers user prompts, while also reducing the volume of data that has to move within the chip. The result is higher inference efficiency.
The name Frozen refers to the idea of “freezing” part of the model’s computational logic into silicon. Google engineers’ 6x to 10x projection refers to token throughput per watt relative to TPU under the same power consumption. The report noted that this figure is still an estimate, not measured performance.
Not a TPU replacement and not planned for mass production
The Information said Google is not planning to manufacture Frozen at the same scale as TPU. The chip is not intended to replace TPU, which serves broader use cases, and Google currently sees part of Frozen v2 as a technical prototype.
Its long-term usefulness also depends on whether Google keeps the same underlying architecture for future Gemini models. If Gemini undergoes major architectural changes, the chip could quickly become outdated. Because the design is tied to a specific architecture, Frozen v2 is not expected to be sold as an external product.
Compute shortage is part of the reason behind the project
The project is partly aimed at easing a serious internal compute shortage at Google. The report said the shortfall has created friction inside the company and has even forced Google Cloud to decline some customer business.
Last month, Google agreed to pay SpaceX nearly $1 billion per month to help fill that compute gap.
Google’s AI business is also facing broader pressure
Beyond compute constraints, Google’s AI push is dealing with other challenges. The report said the release timeline for the next Gemini Pro has slipped, several senior researchers have been poached by rivals, and Chinese models are gradually taking a share of the US enterprise market. New models released over the weekend by Moonshot AI and Alibaba have also drawn market attention.

