Google Delays Gemini 3.5 Pro After Model Falls Short of Internal Targets

Google Delays Gemini 3.5 Pro After Model Falls Short of Internal Targets

N
News Editor
2026-07-17 08:22:09
Google has delayed the release of Gemini 3.5 Pro by months after the model failed to meet internal performance goals, especially in AI coding, according to a Bloomberg report cited in the source material. The model, internally known as “Cappuccino,” had been widely discussed online in the past 48 hours, with leaks pointing to a 2 million-token context window and a new “Deep Think” reasoning mode. Those expectations abruptly reversed after the report said Google had updated training data late last month in an attempt to improve coding performance, only to see disappointing results. The delay quickly spilled into the market. Google shares fell as much as 4.43% after the news. The report also described broader internal problems: complex management layers, competing priorities across major product lines such as Search, Maps, and YouTube, repeated overlap between teams, and limited compute access for engineers using internal AI tools. That tension stands out against the company’s projected 2026 capital expenditure of $180 billion to $190 billion and first-quarter capex of $35.7 billion, more than double a year earlier. The episode has also fed a wider debate about whether frontier AI labs are running into a new scaling wall. Ethan Mollick of Wharton described the pattern as a “next-generation giant-model disappointment trap,” linking Google’s setback to similar concerns around Meta Llama 4 and xAI Grok 4.
GoogleGemini 3.5 ProGoogle DeepMindAI codingLarge language modelsAnthropicOpenAI

Google has pushed back the launch of Gemini 3.5 Pro by months after the model failed to clear internal benchmarks, with AI coding named as the main weakness, according to a Bloomberg report referenced in the source material.

The model, known internally as “Cappuccino,” had been framed as a major response to OpenAI and Anthropic. But the report said development is already months behind schedule, and the latest push to improve the system did not deliver the results Google wanted.

From launch chatter to a sudden delay

In the 48 hours before the report, social platforms were filled with discussion about Gemini 3.5 Pro. Leaks described a 2 million-token context window and a new “Deep Think” mode aimed at stronger reasoning. The model was also said to have improved performance in math, programming, logical reasoning, code generation, agent workflows, frontend UI design, and SVG image generation.

Some online discussion had pointed to July 17 as the expected release date. That changed after Bloomberg reported that Google had updated training data late last month in a last-minute attempt to raise coding performance, but the outcome was described as “disappointing.”

Google Delays Gemini 3.5 Pro After Model Falls Short of Internal Targets 3

After the news broke, Google shares fell, at one point dropping 4.43%.

Coding performance became the sticking point

The report tied the delay directly to coding. As OpenAI and Meta continued to advance new models with stronger code capabilities, Gemini 3.5 Pro failed to hit the level Google expected internally.

Bloomberg also pointed to a deeper cultural issue inside Google. Some veteran engineers have long held the view that important code should be written by humans, not generated by AI. There were also concerns about proprietary code leaking into training data, which limited the use of Gemini in internal development workflows.

Google Delays Gemini 3.5 Pro After Model Falls Short of Internal Targets 4

When Google later moved to require broader use of AI for coding, another problem surfaced. Engineers trying to use internal AI tools repeatedly ran into compute limits.

Heavy spending did not remove compute bottlenecks

That mismatch stood out in the report. Google is projected to spend $180 billion to $190 billion in capital expenditures this year. Wall Street data cited in the source also said the company spent $35.7 billion in the first quarter, more than double the level from a year earlier.

Even with that level of investment in chips and data centers, internal engineers were still running into GPU shortages when using the company’s own AI systems.

Google Delays Gemini 3.5 Pro After Model Falls Short of Internal Targets 5

Google is now trying to consolidate its internal AI coding stack. The report said the company’s chief AI architect is moving separate coding tools onto the Google Antigravity foundation, while DeepMind has created a dedicated AI coding team.

Coordination across large divisions slowed execution

Bloomberg’s account described a broader structural problem inside Google. Launching a model of this scale requires coordination across major businesses including Search, Maps, and YouTube, which brings in multiple layers of management and competing priorities.

A former employee quoted in the report compared the process to “trying to boil the ocean” if the goal is to get every leadership team pulling in the same direction. The result, according to the report, was shifting instructions, duplicated work across departments, and slower decisions.

Google Delays Gemini 3.5 Pro After Model Falls Short of Internal Targets 6

Google has several major groups involved in AI, including Google DeepMind, Google Cloud, and Android teams, and it has also formed multiple internal units to work on AI coding. The report suggested that this internal “horse race” approach showed commitment, but also created overlap and friction.

Researchers were said to be leaving for rivals

The same tensions have weighed on retention. According to the report, many researchers frustrated by Google’s slower progress have left for Anthropic and OpenAI.

That sets up a damaging loop: bureaucracy drags on execution, weaker execution leads to product slippage, product slippage feeds talent departures, and talent losses make it harder to catch up. In the source material, the delay of Gemini 3.5 Pro was presented as a clear example of that cycle.

Google Delays Gemini 3.5 Pro After Model Falls Short of Internal Targets 7

A wider debate over the next wave of frontier models

Ethan Mollick of the Wharton School used the report to raise a broader point. In a post on X referenced by the source, he argued that Google’s setback may not be an isolated case and described the pattern as a “next-generation giant-model disappointment trap.” He linked it to the difficulties seen with Meta Llama 4 and xAI Grok 4.

The idea is that huge spending on the next generation of models is no longer producing gains that match expectations, putting pressure on market leadership.

The source listed several pressures behind that view: high-quality human text data is close to exhaustion, the value of synthetic data remains uncertain, Transformer-based architectures may be approaching their limits, and marginal performance gains now require sharply rising compute costs.

Google Delays Gemini 3.5 Pro After Model Falls Short of Internal Targets 8

The same material said OpenAI, with Orion/GPT-4.5, has so far avoided a major slide under those conditions.

Gemini 3.5 Pro’s delay has sharpened attention on how hard it is becoming to keep pushing frontier models forward. As model scale moves closer to engineering and physical constraints, each new step appears harder to achieve.

The cited references include posts on X from Mr_Salio and Ethan Mollick, a Bloomberg article dated July 16, 2026, and the original Chinese source, which credited the WeChat account New Intelligence with authorship attributed to “ASI启示录.”

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.