Google has pushed back the launch of Gemini 3.5 Pro by months after the model failed to clear internal benchmarks, with AI coding named as the main weakness, according to a Bloomberg report referenced in the source material.
The model, known internally as “Cappuccino,” had been framed as a major response to OpenAI and Anthropic. But the report said development is already months behind schedule, and the latest push to improve the system did not deliver the results Google wanted.
From launch chatter to a sudden delay
In the 48 hours before the report, social platforms were filled with discussion about Gemini 3.5 Pro. Leaks described a 2 million-token context window and a new “Deep Think” mode aimed at stronger reasoning. The model was also said to have improved performance in math, programming, logical reasoning, code generation, agent workflows, frontend UI design, and SVG image generation.
Some online discussion had pointed to July 17 as the expected release date. That changed after Bloomberg reported that Google had updated training data late last month in a last-minute attempt to raise coding performance, but the outcome was described as “disappointing.”

After the news broke, Google shares fell, at one point dropping 4.43%.
Coding performance became the sticking point
The report tied the delay directly to coding. As OpenAI and Meta continued to advance new models with stronger code capabilities, Gemini 3.5 Pro failed to hit the level Google expected internally.
Bloomberg also pointed to a deeper cultural issue inside Google. Some veteran engineers have long held the view that important code should be written by humans, not generated by AI. There were also concerns about proprietary code leaking into training data, which limited the use of Gemini in internal development workflows.

When Google later moved to require broader use of AI for coding, another problem surfaced. Engineers trying to use internal AI tools repeatedly ran into compute limits.
Heavy spending did not remove compute bottlenecks
That mismatch stood out in the report. Google is projected to spend $180 billion to $190 billion in capital expenditures this year. Wall Street data cited in the source also said the company spent $35.7 billion in the first quarter, more than double the level from a year earlier.
Even with that level of investment in chips and data centers, internal engineers were still running into GPU shortages when using the company’s own AI systems.

Google is now trying to consolidate its internal AI coding stack. The report said the company’s chief AI architect is moving separate coding tools onto the Google Antigravity foundation, while DeepMind has created a dedicated AI coding team.
Coordination across large divisions slowed execution
Bloomberg’s account described a broader structural problem inside Google. Launching a model of this scale requires coordination across major businesses including Search, Maps, and YouTube, which brings in multiple layers of management and competing priorities.
A former employee quoted in the report compared the process to “trying to boil the ocean” if the goal is to get every leadership team pulling in the same direction. The result, according to the report, was shifting instructions, duplicated work across departments, and slower decisions.

Google has several major groups involved in AI, including Google DeepMind, Google Cloud, and Android teams, and it has also formed multiple internal units to work on AI coding. The report suggested that this internal “horse race” approach showed commitment, but also created overlap and friction.
Researchers were said to be leaving for rivals
The same tensions have weighed on retention. According to the report, many researchers frustrated by Google’s slower progress have left for Anthropic and OpenAI.
That sets up a damaging loop: bureaucracy drags on execution, weaker execution leads to product slippage, product slippage feeds talent departures, and talent losses make it harder to catch up. In the source material, the delay of Gemini 3.5 Pro was presented as a clear example of that cycle.

A wider debate over the next wave of frontier models
Ethan Mollick of the Wharton School used the report to raise a broader point. In a post on X referenced by the source, he argued that Google’s setback may not be an isolated case and described the pattern as a “next-generation giant-model disappointment trap.” He linked it to the difficulties seen with Meta Llama 4 and xAI Grok 4.
The idea is that huge spending on the next generation of models is no longer producing gains that match expectations, putting pressure on market leadership.
The source listed several pressures behind that view: high-quality human text data is close to exhaustion, the value of synthetic data remains uncertain, Transformer-based architectures may be approaching their limits, and marginal performance gains now require sharply rising compute costs.

The same material said OpenAI, with Orion/GPT-4.5, has so far avoided a major slide under those conditions.
Gemini 3.5 Pro’s delay has sharpened attention on how hard it is becoming to keep pushing frontier models forward. As model scale moves closer to engineering and physical constraints, each new step appears harder to achieve.
The cited references include posts on X from Mr_Salio and Ethan Mollick, a Bloomberg article dated July 16, 2026, and the original Chinese source, which credited the WeChat account New Intelligence with authorship attributed to “ASI启示录.”

