Google on July 21 unveiled a new set of Flash-series models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. In a post on the company’s official blog, Tulsee Doshi, Senior Director of Product Management, said the models were built for AI agent workflows and aimed at large-scale enterprise deployment, with a focus on lower output costs, lower token consumption and stronger computer-use capabilities.
The announcement comes as AI models are increasingly being judged on how well they balance speed, cost and reasoning for agentic use cases. Google said the three releases were shaped by prior user feedback and are meant to strengthen its infrastructure position in the AI agent market.
Gemini 3.6 Flash targets coding, multimodal work and lower token usage
Google described Gemini 3.6 Flash as its new primary work model. The company said it posted stronger results in coding, knowledge work and multimodal tasks. On OSWorld-Verified, a benchmark for computer-use capability, Google said the model scored 83.0%. The related computer-use tools are now built directly into the Gemini API and enterprise offering.
Google also highlighted token efficiency. Citing the Artificial Analysis Index, the company said output token usage for Gemini 3.6 Flash is down 17% from the previous generation. On some benchmarks, including DeepSWE, the reduction reached as high as 65%.
Google said those changes cut reasoning time and tool-calling time in multi-step workflows while making pricing more competitive. The listed price is $1.50 per million input tokens and $7.50 per million output tokens.
Gemini 3.5 Flash-Lite is built for speed-sensitive workloads
For tasks that need low latency and high throughput, such as agentic search and document processing, Google introduced Gemini 3.5 Flash-Lite. According to the company, the model can generate up to 350 tokens per second. Pricing was listed at $0.3 per million input tokens and $2.5 per million output tokens.
Google said the model, while lightweight, delivers a quality level that exceeds earlier Lite generations. It also supports adjustable “thinking levels,” letting developers switch between a lower-thinking mode and a higher-thinking mode for more complex sub-agent workflows. In long-context and agentic coding evaluations, Google said Flash-Lite outperformed the standard Gemini 3 Flash model.
Cybersecurity model arrives as Gemini 4 pre-training begins
Google also released Gemini 3.5 Flash Cyber, a fine-tuned version of 3.5 Flash designed for specific cybersecurity defense scenarios. The company said the model can work with the multi-agent system CodeMender to produce consolidated reports that help defenders identify and fix security vulnerabilities more quickly.
Access to Gemini 3.5 Flash Cyber is limited for now. Google said the model is available only to government agencies and trusted partners through a pilot program because of the dual-use risks tied to this type of technology.
At the end of the post, Doshi said Google is testing Gemini 3.5 Pro with partners ahead of a broader rollout. She also said the company has formally started pre-training Gemini 4, calling it Google’s “most ambitious” project. Developers can now access Gemini 3.6 Flash and Gemini 3.5 Flash-Lite through the Gemini API and other channels.

