Google on Oct. 6 introduced Nano Banana 2.1, an image generation model the company says improves on its predecessor across visual design, mask-based editing and subject consistency.
The launch came with a notable price cut on image output. According to Google’s pricing page, 1K, 2K and 4K images now cost $0.0336, $0.0504 and $0.0756 each through the API. The prior model was priced at $0.067, $0.101 and $0.151 for the same sizes. For 1,000 1K images, the bill comes to about $33.6, down from $67.
That reduction applies only to image output. Google also raised other pricing components. Input costs per 1 million tokens increased from $0.50 to $1.50, while text and thinking output rose from $3 to $7.50. Workflows that pack in more reference images or use higher thinking levels could see part of the image-price reduction offset by those higher token charges.
Google’s benchmarks put Nano Banana 2.1 ahead across all 10 tests
The model card lists 10 evaluations. With thinking enabled, Nano Banana 2.1 leads both Nano Banana 2 and Nano Banana Pro in every one of them. Google describes thinking as the model’s internal drafting effort before generating an image, with three levels: minimal, medium and high. Medium is the default.
On overall preference, Nano Banana 2.1 posted an Elo score of 1050, compared with 990 for the previous model and 935 for Pro. In this setup, human raters compare two outputs side by side, pick the one they prefer, and the results are converted into Elo ratings. A higher score means the model’s output was chosen more often.
For multi-character consistency, Nano Banana 2.1 scored 1106, versus 978 for Nano Banana 2 and 1011 for Pro. For mask editing, it scored 1049, compared with 965 and 927.
Factuality in infographics was measured by an automated grader rather than human preference tests, and the model card does not explain the scale. Nano Banana 2.1 scored 0.521. The earlier model scored 0.179, and Pro scored 0.265.
With thinking turned off, Nano Banana 2.1 still reached 1015 in overall preference, above the predecessor’s 990 even with thinking on. Google’s own numbers suggest the gains are not solely the result of an extra reasoning step.
Those results come from Google-run evaluations, and the comparison table includes only Google models. The company’s release does not cite third-party validation.
Up to 14 reference images, plus fixes for extreme aspect ratios
On the specification side, Nano Banana 2.1 supports up to 14 reference images. For character consistency, it supports as many as 4 characters. For object reconstruction, the limit is 10 objects.
Google also says it fixed stitching artifacts that appeared at 2K and 4K resolutions with very tall or very wide aspect ratios such as 1:4 and 8:1.
The model card lists multiple access points besides the Gemini API: Gemini App, Google Search AI Mode, Google AI Studio, Google Ads, Flow and Stitch.
Four official examples highlight text rendering, masking and consistency
Google’s DeepMind showcase page includes grouped image examples and the prompts used to produce them.
Visual design and text layout
The first example pairs two posters. On the left is a desert trail-running poster labeled “DESERT,” filled with small text including HEAT TOLERANCE, ZERO HORIZON, GPS coordinates, ELEV. 1,850 M and a barcode. On the right is a 1970s-style travel poster. The prompt asks for the large word “HORIZON” to be partially obscured by a woman sitting on top of a vintage camper van. Google highlights that the requested text is spelled correctly and that the overlap between the person and typography creates depth.
The official prompt reads: 「Cinematic 1970s travel poster. A thin warm cream border frames the entire composition. Along the top header border is tiny, spaced, uppercase black text reading: “PACIFIC COAST ISSUE 08 NORTHERN SWELLS HIGHWAY ONE”. The background is a solid flat field of vibrant cobalt blue. Across the upper half, giant ultra-condensed cream-colored sans-serif text spells “HORIZON”. In the foreground, a young woman in a denim vest and sunglasses sits atop a vintage seafoam-teal camper van with surfboards mounted on top, overlapping and partially masking the giant letters to create depth. Rich film grain texture, warm natural lighting, high-contrast, ultra-sharp legible typography, no placeholder text, all text in English.」
Mask-based editing
The second example shows mask editing. The input image is a dew-covered dandelion seed circled by a hand-drawn sketch. The outputs place that seed into four different settings: a mossy tree trunk, autumn leaves, a beach and a tiled tabletop. The prompt is a single sentence asking the model to keep the object’s size and position unchanged while replacing only the environment. The original request was for five images, though the showcase displays four.
The official prompt reads: 「mask the object inside the sketch and place in a completely different environment but don’t change the size and position of the object itself, create 5 of these each 16:9, each it’s own server call, one at a time」
Google’s point here is that the seed stays nearly fixed in size and placement while the surrounding scene changes.
Subject consistency
The third example focuses on subject consistency. The left side contains 14 sets of fashion model reference images, each set showing the same person from multiple angles. The right side combines them into one editorial-style group shot inside a geometric, color-blocked architectural setting, with some subjects placed closer to the viewer and others farther away. The prompt is only a short, conversational instruction and leaves the theme to the model.
The official prompt reads: 「i want a high-end editorial fashion shot with all these models posing together in abstract colorful set, dramatic composition, some models are far and some are close. choose a theme」
The showcase emphasizes that multiple outfits need to match their references while preserving depth in a shared scene.
Image quality and natural feel
The fourth example lines up with Google’s claim of “more natural-looking images.” It shows a woman in an emerald-green mini dress descending a mustard-yellow exterior staircase at a retro motel, set against a flamingo-pink stucco wall. The prompt describes the composition, location and yellow-pink-green color contrast, but does not explicitly ask for shadow detail or weathered textures.
The official prompt reads: 「A 16:9 horizontal cinematic still, shot from street level looking up at the sharp diagonal steel exterior staircase of a two-story retro motel painted vibrant mustard-yellow against a flamingo-pink stucco facade. An adult woman in an emerald-green mini dress descends the staircase, her hand trailing along the yellow railing. The bright yellow, pink, and green color clash creates a delightful, sun-drenched cinematic mood.」
Google points to details the model filled in on its own, including the railing shadow cast on the wall and the worn texture of the stucco surface.
Previous model set to shut down on Oct. 29
Google set a retirement date for the previous production model the same day it launched Nano Banana 2.1. The production version, gemini-3.1-flash-image, went live on May 28 and is scheduled to shut down on Oct. 29, giving it a lifespan of roughly five months. Google recommends that developers move to gemini-nano-banana-2.1.
From Oct. 6 to Oct. 29, the migration window is 23 days.
This is also the first time the Nano Banana series has used the nickname directly in the model string. Earlier versions used formal naming only.
Compatibility notes and listed limitations
Developers planning a switch also need to check supported sizes. Nano Banana 2.1 does not support the 512px minimum size. Even so, the previous model’s 0.5K image cost $0.045, while a 1K image in 2.1 costs $0.0336.
The model card also lists several limitations. Small text at 1K resolution is often blurry. Character consistency is not perfect every time. The model can occasionally confuse left and right, and it may slow down or time out.

