Google is reportedly developing a new AI chip meant for one job only: running Gemini faster and more cheaply.

The Information reported Monday that the chip is internally known as Frozen v2. The project gives Google a possible answer to a problem money alone has not fixed quickly enough: the company is running short on capacity to serve the AI demand it has already created.
In March, Google told Meta it could not meet the amount of Gemini compute Meta wanted to purchase, according to the report. Meta then had to tell employees to ration their AI use. Google, which is spending as much as $190 billion on AI infrastructure this year, was also said to be turning customers away because it did not have enough servers.
How Frozen v2 differs from Google’s TPU line
Only limited details about the new chip have been reported so far. By naming convention, though, Frozen v2 does not appear to be another routine upgrade to Google’s Tensor Processing Units, or TPUs, the custom chips Google has been building since 2015 to power Gemini and its cloud services for outside developers.
Those Tensor chips can run any AI model loaded onto them. Frozen v2, according to the report, is designed differently. It would bake part of Gemini’s architecture directly into the hardware. That architecture is the structural blueprint that determines how the model routes and processes information.
In machine learning, “freezing” usually means locking something in place permanently. In this case, what would be frozen is the architecture rather than the model’s weights, the knowledge Gemini picks up in training and can still update later.
By hardwiring that blueprint into the chip’s circuits, Frozen v2 could skip some redundant calculations and avoid moving data across memory on every query. Engineers reportedly project a 6x to 10x improvement in the number of tokens generated per watt of electricity consumed. Tokens are the small text units that make up each AI response.
As framed in the report, that is the difference between serving ten queries for the power cost of one.
Lower running costs matter more than user-facing changes
If deployed, the chip would not necessarily change how Gemini feels to end users. The bigger shift would be on the cost side. A Gemini model that is cheaper to run would compete more aggressively with OpenAI, Anthropic, and Chinese AI labs, which already account for as much as 45% of U.S. company AI token usage, largely because they operate at costs that are 60% to 90% lower, the report said.
That does not mean users would immediately get cheaper AI products. It could, however, improve Google’s own profitability.

Stock reaction before earnings
Alphabet shares rose roughly 3% during Monday’s trading session on the report and touched $356 intraday. The move faded in the following session as investors turned their attention to Google’s latest results.
The company is scheduled to report Q2 2026 earnings on Wednesday, July 22.
Another push to reduce Nvidia dependence
The reported Frozen v2 effort is also part of a broader move by major AI companies to reduce dependence on Nvidia hardware. Nvidia controls about 85% of the AI GPU market, and large technology companies have been trying to build alternatives.
According to the article, Nvidia’s hardware was originally built for video games rather than language models. It works for AI workloads, but it carries overhead that purpose-built chips can avoid. At Google’s scale, a 6x to 10x efficiency gap translates into real money, potentially billions of dollars.
That is why Meta, Amazon, Microsoft, and OpenAI are all pursuing custom silicon programs of their own. Decrypt noted that even Amazon Web Services, despite committing to deploy 1 million Nvidia GPUs through 2027, is developing its own chips at the same time to reduce long-term exposure.
Still exploratory, with deployment no earlier than 2028
Frozen v2 remains an exploratory project, according to the report. Key design decisions have not been finalized, Google has not confirmed that the effort exists, and the chip would not be offered to outside cloud customers because hardware wired for one model cannot run someone else’s.
Deployment is targeted for 2028 at the earliest, the report said.
Until then, Google is relying on interim capacity. The article says Google is paying SpaceX $920 million a month to rent 110,000 Nvidia GPUs from xAI data centers as a bridge.

