Google released Nano Banana 2.1, an image generation and editing model based on Gemini 3.6 Flash, on October 6, 2026, with availability across the Gemini app, Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform and paid-tier API image output priced at $30 per million tokens.
The Nano Banana 2.1 model card, published the same day, describes the model as a member of the Gemini 3 series of natively multimodal reasoning models. It accepts text strings and images as input with a context window of up to 1 million tokens, and it generates image and text output, listed with 4K token and 64K token outputs respectively. The card lists distribution through the Gemini app, Google AI Studio, the Gemini API, AI Mode in Google Search, Google Ads, Google Flow, and Google Stitch.
Enterprise Platform Specifications
On the Gemini Enterprise Agent Platform, the model carries the identifier gemini-nano-banana-2.1 at the generally available launch stage, with a listed release date of October 6, 2026. Google’s platform documentation describes Nano Banana 2.1 as optimized for multimodal image generation and editing, offering a balance of price and performance.
The documentation lists text and image as supported for input and output, video as input-only, and audio as unavailable, with a 131,072-token context window and a maximum of 32,768 output tokens. The seed, topK, logprobs, temperature, and topP parameters are not supported, and setting any of them returns an API error.
Supported features include reasoning, system instructions, implicit context caching, token counting, grounding with Google Search and image search, provisioned throughput, batch inference, and standard pay-as-you-go, with global availability. Image generation, including from video inputs, image editing and multi-turn editing, interleaved image and text, Content Credentials (C2PA), and virtual try-on are supported; people generation is unavailable.
Limits include 14 images per prompt, 7MB per file for inline data or console uploads, 30MB per file from Cloud Storage, and a 500MB input size limit. Supported aspect ratios include 1:1, 3:2, 4:3, 16:9, 9:16, and 21:9, with 1K, 2K, and 4K resolutions and png, jpeg, webp, heic, and heif file support. The documentation states the model consumes 1,120 tokens per input image, with output images consuming 1,120 tokens at 1K resolution and 1,680 tokens at 2K.
Gemini API Pricing
Google’s Gemini API pricing page, updated October 6, 2026, describes Nano Banana 2.1 as “an update to Nano Banana 2 (Gemini 3.1 Flash Image),” built for high-efficiency image generation and conversational editing with improved visual quality, multi-turn character consistency, accurate text rendering, and search-grounded generation across 1K, 2K, and 4K resolutions. Standard paid-tier pricing is listed at $1.50 per million input tokens for text, image, and video, $7.50 per million output tokens for text and thinking, and $30.00 per million tokens for image output, which the page equates to $0.0336 per 1K image, $0.0504 per 2K image, and $0.0756 per 4K image.
Batch pricing is listed at $0.75 per million input tokens, $3.75 per million tokens for text and thinking output, and $15.00 per million image tokens, with stated equivalences of $0.0168, $0.0252, and $0.0378 per 1K, 2K, and 4K image respectively. Grounding with Google web and image search carries 5,000 free requests per month shared across Gemini 3.x models, then $14 per 1,000 requests. The page lists no free-tier access for Nano Banana 2.1.
The same page lists the predecessor, Gemini 3.1 Flash Image (Nano Banana 2), at $0.50 per million input tokens, $3 per million tokens for text and thinking output, and $60.00 per million image tokens, with per-image equivalences of $0.067 at 1K, $0.101 at 2K, and $0.151 at 4K.
Reported Evaluations
The model card reports an evaluation approach combining side-by-side human evaluation producing Elo scores across text-to-image and editing tasks, a single-sided AutoRater for factuality, and regression sets drawn from Gemini 2.5 Flash Image and Gemini 3 Pro Image use cases.
For text-to-image, the card reports an Overall Preference Elo of 1050 ±14 for Nano Banana 2.1 in its Thinking configuration and 1015 ±13 without thinking, against 990 ±7 for Nano Banana 2 (Thinking) and 935 ±8 for Gemini 3 Pro Image (Nano Banana Pro). Reported Infographic Design scores are 1048 ±17 for the Thinking configuration, against 961 ±12 for Nano Banana 2 and 912 ±12 for Nano Banana Pro, while Infographic Factuality is listed at 0.521, against 0.179 and 0.265.
For editing in the Thinking configuration, the card reports Nano Banana 2.1 at 1026 for General Editing, 1028 for Single-Character Consistency, 1106 for Multi-Character Consistency, 1049 for Mask/Ink-Based Editing, 1024 for Product Consistency, 1062 for Stylization, and 1066 for Multi-Reference Editing, each listed above the corresponding Nano Banana 2 and Nano Banana Pro scores in the same table.
Limitations and Safety Assessments
The card lists known limitations including hallucinations, occasional slowness or timeouts, poor rendering of small text and long paragraphs, imperfect character consistency between input and generated images, partial instruction following and ink persistence in masked editing, rare instances of persistent subject pose during editing, occasional confusion around spatial localization, and limited world knowledge, 3D reasoning, and factuality. Gemini 3.6 Flash has a knowledge cutoff of March 2026, though the card notes that in some domains knowledge may be limited to January 2025.
The card reports manual red teaming by specialist teams outside the model development team, states that Nano Banana 2.1 satisfied required child-safety launch thresholds, and describes its safety performance as similar or improved compared with Gemini 3 Flash, with no egregious concerns found relative to Gemini 3.1 Pro. For frontier safety, the card states that Gemini 3.1 Pro and Gemini 3.7 Flash did not reach any Tracked or Critical Capability Levels and that Nano Banana 2.1 shows no meaningful new capabilities or material performance increases over those models, leaving it assessed as unlikely to reach any such levels.
Credit: Source link



























