Gemini 4 Pro has surfaced again through a new round of leaked checkpoints, with developers identifying two fresh versions under the internal label "barium-b" or Checkpoint 2. The demos now circulating online point to sharp gains in 3D physical simulation, code generation and long-context reasoning.

One developer quoted in the source said Gemini 4 Pro still felt startling even after spending 400 hours using Opus 5.5 and GPT-6. The article frames Google’s next-generation model as especially strong in 3D physics, coding and long-form reasoning tasks.
Public demos focus on 3D generation, physics engines and long-form code building
According to the report, developers quickly began sharing demos after the new checkpoints were found. The most discussed examples center on 3D generation, physical simulation and large code outputs. The article also mentions a peacock SVG test, saying it showed a major jump from earlier results.
Floatplane takeoff test against Opus 5.5
Reviewer @AI_Screening asked Gemini 4 Pro and Opus 5.5 to generate, through code, a 3D physics simulation of a floatplane taxiing on water and taking off. In the comparison described by the source, Opus 5.5 produced a moving aircraft, but the fuselage shook heavily and the water response looked stiff. Gemini 4 Pro’s version was described as smoother, with steadier acceleration, clearer reflections on the water and more convincing splash effects.

A 3D International Space Station in 16 minutes
Developer @thtbee_ posted another test. Without importing any 3D model, Gemini 4 Pro reportedly used pure code to build an interactive simulation of the International Space Station orbiting Earth in 16 minutes. The article says it created the station model with Three.js and pulled the correct Earth texture from the web so that latitude, longitude and color would match properly.
The source also notes shortcomings. The generated UI was criticized as unattractive, and the back side of the model still showed a perspective bug.
From lighthouse to rocket launch
User @HarshithLucky3 gave the model a harder instruction: "Use Three.js to turn a lighthouse into a rocket launch, without any UI elements." The article says Gemini 4 Pro followed that instruction cleanly. The tester highlighted the texture quality across the scene and praised the overall 3D output.
High-detail 3D Wii remote
Developer @LuminaBench tested industrial design modeling by asking for a 3D Nintendo Wii remote. The result, according to the source, showed strong control over object detail and unusually fast generation speed. At the same time, the article says Gemini 4 Pro sometimes overworked prompts, adding extra detail even when it was not requested.

The same feedback said its design taste was not equally strong across every component. Some parts looked excellent, while others felt mismatched. The report adds that Checkpoint 2 stood out for detail density, higher precision and very fast output speed.
Mechanical hummingbird benchmark
Blogger @TimJayas used a long-running mechanical hummingbird prompt to test the model’s grasp of complex mechanical structures and dynamic behavior. The conclusion cited in the article was that Gemini 4 Pro performed extremely well, with output quality on par with Fable and well ahead of GPT-6 Sol. The same passage adds that GPT-6 Astra still beats it in some cases.
Full website in 14 minutes, games from a single prompt
The report also stresses engineering output. One developer, it says, generated a complete product showcase website in 14 minutes, covering both front end and back end. Multiple users also said a single prompt was enough for Gemini 4 Pro to produce a fully interactive web game.
Leaked specs mention a 10 million-token context window and 256k output
Although the new checkpoints only surfaced recently, the article says some Gemini 4 Pro specifications had already leaked. Those claims include a 10 million-token context window and a 256k output limit.

If accurate, the source argues, those numbers would reduce the problem of long code generations cutting off midway. The same leak claims persistent memory across sessions, web access without an API, agent capabilities and even robotics control, with the model acting as the "brain" for machines in the physical world.
A benchmark image referenced in the article also said Gemini 4 Pro outperformed GPT-6 Astra and Claude Fable 5.1 in coding, agent tasks and reasoning. It reportedly scored 88% on the AI agent programming test DeepSWE v1.1 and 95.3% on the coding benchmark Terminal-bench 2.1.
Google DeepMind chief says Gemini 4 is already in early post-training
The article links the sudden appearance of Gemini 4 Pro to a shift in Google’s internal development pace.
According to The Information, Google DeepMind chief Koray Kavukcuoglu publicly confirmed for the first time that Gemini 4 moved into post-training after what he described as Google’s most ambitious pretraining effort ever. That effort began in July, and early pretraining was largely completed in about two months, the report says.

"Our intention is to release an early post-trained version as quickly as possible," Koray said. He added, "Because we’ve seen the results, we are very excited."
The article argues that this helps explain why a checkpoint that still appears to be an early build has already shown up in public testing arenas.
The report also says Google shifted its compute resources to Gemini 4 after deciding Gemini 3.5 Pro was not strong enough to dominate rivals. When asked whether Google had fallen behind competitors, Koray replied: "I have 100% confidence in my team. We are destined to always be at the forefront of technology."
Internal use, TPU design and launch timing
Koray also said Gemini 4 is being used internally by Google engineers for Antigravity AI. The article presents that as a response to earlier rumors that Google employees were using Claude for coding work.

On the broader goal, he said the company is no longer fixated on the AGI label and is instead asking whether it can build agents that can be fully trusted.
He also said Gemini 4 development has helped Google hardware engineers design the next two to three generations of TPU chips.
As for timing, Koray said he hopes Gemini 4 will launch "well before the end of the year." The article separately says there is also talk that Gemini 4 Pro could arrive around China’s National Day holiday period.
Leaked pricing and a satellite mission carrying TPU chips
The report includes a leaked pricing structure for Gemini 4 Pro:

- Input: $2.25 per 1 million tokens
- Output: $11.25 per 1 million tokens
The article describes that pricing as highly aggressive.
It also says Google will send a satellite carrying four TPU chips into space next week aboard a SpaceX Falcon 9 rocket. The source then points to the possibility of a future satellite-based AI data center network.
The piece closes by quoting a tester who said the second checkpoint showed a "visible leap" over the first. It leaves open the broader question of how OpenAI and Anthropic will respond to Gemini 4 Pro’s arrival.

