Google briefly added its Nano Banana 2 image model to the web version of Google Earth, then pulled the feature less than a day later. The source article says the company removed it after highly realistic outputs drew strong objections from experts, and plans to release it again after adding stronger safeguards.
The update placed a new “Create image” button in the Google Earth web interface. Instead of exporting screenshots to other AI tools, users could stay inside the browser, pick a location, enter a prompt, and generate a new visual directly on top of the selected geographic scene.
How Google positioned the feature
The article listed several examples presented by Google. One was historical reconstruction. A teacher could open a historical site and enter a prompt asking for a photorealistic image showing what Pompeii looked like in 78 AD. The result would turn present-day ruins into a vivid Roman streetscape for classroom use.
Another use case was landmark infographics. A user could ask for an easy-to-understand infographic of the Statue of Liberty labeled with key historical facts. According to the article, Gemini would retrieve the background information, while Nano Banana would generate the visual in a matter of seconds.

The tool was also framed as a way to preview development concepts. Architects and urban planners could select an empty site in Tokyo and ask the model to reimagine it as an active shopping district with open public space. The generated image would replace the bare lot with a polished 3D-style urban rendering.
For smaller-scale planning, users could zoom in on an undeveloped parcel and prompt the system to add a modern lakeside cabin built with local sustainable materials. The article said the system could then return a photo-real preview integrated with the real terrain.
It also supported more imaginative experiments. One example described turning Google’s Mountain View campus into a futuristic utopia with glowing walkways, glass eco-domes, flying transport pods, and trees growing around the buildings, transforming a real place into a cinematic sci-fi city.
A tester turned Philadelphia into a dystopian scene
Technology editor David Gewirt quickly tried the feature after learning about it. The article said he had already spent hours exploring Google Earth in the past and remained fascinated by what it allowed users to see from home.

He used Independence Hall in the United States as one test case. The building is a U.S. National Historic Landmark and a World Heritage Site, and the Declaration of Independence was adopted there on July 4, 1776. His prompt asked the model to imagine the site in a dystopian future 500 years from now. The result turned the landmark into a ruined post-apocalyptic structure.
He then pushed a Philadelphia Mummers Parade scene further. The article described the parade as the oldest annual folk parade in the United States and a New Year tradition since 1901. Gewirt asked the model to have the area “occupied by zombies, evil clowns, and giant alien mechs,” then added another instruction telling the AI to make them all happy.
He said the results “caused a lot of laughs” and called the output “a plausible future Philadelphia, especially on January 1.” Even so, he added a caution: “While this can lead to some fun experiments, it’s clearly not yet a reliable, practical tool.”
The source article also noted that the feature does not currently work directly in Street View. That limitation cuts into both its usefulness and its entertainment value. One media outlet was even harsher, saying that for professional developers or architects, imagery generated without guaranteed geometric rigor and assembled through probabilistic methods is just polished “AI slop.”

What changed technically: geospatial grounding
The article described the key shift as a move from unconstrained image generation to “geospatial grounding.” In earlier AI image systems, a model could assemble a picture from prompt-based probabilities without needing to respect gravity, terrain, or whether a building actually exists in the real world. Inside Google Earth, Nano Banana 2 is placed under a much stricter set of conditions.
It said the setup relies on Gemini’s multimodal and search-grounding capabilities and works through at least three layers.
The first is live capture of the physical base map. Nano Banana 2 is not working from a simple top-down image. Its inputs are described as a compound constraint matrix made up of the current satellite or aerial view, 3D elevation and terrain meshes, and spatial camera parameters.
Under that setup, generation is anchored to real geography. Terrain cannot be changed freely. Mountain folds, lake edges, and the position and scale of building foundations have to remain consistent with actual conditions. Newly generated buildings, vegetation, and scene elements also have to fit the existing 3D terrain and structures rather than float in midair.

The article added that Nano Banana 2 supports consistent rendering for as many as five subjects and 14 objects, which is meant to reduce visual collapse when the viewpoint changes. This is framed as conditioned generation under real geographic constraints. The trade-off is speed: each image in Google Earth can take as long as two minutes to produce.
The second layer is live retrieval from a world knowledge base. The article said Nano Banana 2 is connected to Google Search and Google’s geographic knowledge graph. When a user tries to reconstruct history or annotate a landmark, the system can retrieve cultural background and geographic context tied to that place.
For the Statue of Liberty infographic example, the model does not only draw the image. It also calls on Gemini’s world knowledge to fill in construction year, materials, dimensions, and historical information. The article also made a point of warning that retrieval is not automatically perfect. An infographic that looks professional does not guarantee that every fact in it is correct.
The third layer is high-fidelity output and stronger text rendering. According to the article, Nano Banana 2 can generate 2K and even 4K images and is able to render on-image text much more clearly than older systems. That allows users to create scenes with legible signs, labels, and educational overlays placed directly on top of landmarks.

Workflow and project saving
On the product side, the workflow was simplified heavily. During generation, users could step forward and backward to review changes, then click “Refine image” to add more prompts and keep iterating. The interface also offered a clear before-and-after view.
Once satisfied, a user could click “Save to Project.” The generated image would then be pinned to a personal project layer as a Placemark while preserving the camera angle. Anyone opening that coordinate could see the alternative version of the place.
Why the article says this matters beyond a single feature
The source argued that bringing Nano Banana into Google Earth could mean more than shipping another image generator.
It placed the move inside a crowded AI image market. OpenAI’s GPT-image-2 was said to be leading the Arena rankings, while Nano Banana 2 Lite ranked fifth. Adobe Firefly and Midjourney were also named as rivals competing for the same users.

Google’s advantage, the article argued, is not just model architecture but data. It pointed to two decades of satellite imagery and aerial photography, along with 3D terrain models covering hundreds of cities across more than 40 countries. Combined with Nano Banana 2, that proprietary spatial dataset creates a new category of AI visualization tied to real coordinates and spatial credibility.
In that framing, the contest is no longer about which model makes the prettiest image. It becomes a contest over which one can make the most believable image, down to latitude, longitude, elevation, and physical viewpoint. The article closed by invoking Kevin Kelly’s idea of a “mirror world,” arguing that Google is layering a prompt-editable digital skin on top of that mirrored version of the physical world.
The references listed in the article include Google’s official blog, ZDNET, TechTimes, and X posts from Google Earth and Google. The piece was published by MarsBit, credited to the WeChat account New Intelligence with the author name “ASI启示录,” and edited by David.

