New Gemini Models Target Agentic Workflows and Efficiency
Google DeepMind has released three new models in its Gemini Flash series, aiming to give developers and enterprises better tools for building AI agents at scale. The lineup includes Gemini 3.6 Flash, 3.5 Flash-Lite, and a specialized cybersecurity model called 3.5 Flash Cyber. The company also noted that Gemini 3.5 Pro is currently being tested with partners and will be made broadly available when ready. Meanwhile, work has begun on the next generation, with DeepMind starting its most ambitious pre-training run yet for Gemini 4.
Tulsee Doshi, Senior Director of Product Management at Google DeepMind, announced the releases on behalf of the Gemini team. The new models are designed to address the needs of developers and customers who require higher token efficiency, lower latency, and more reliable performance for production AI agents. The Flash series has been built to balance efficiency and quality, enabling the scaling of agentic workflows.
Gemini 3.6 Flash: Better Performance, Lower Cost
Gemini 3.6 Flash is positioned as a workhorse model that improves on its predecessor, 3.5 Flash, in coding, knowledge work, and multimodal tasks. According to the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash. In some benchmarks, such as DeepSWE by Datacurve, the reduction in output token usage reaches up to 65%. The model also takes fewer reasoning steps and tool calls to complete multi-step workflows.
Pricing for 3.6 Flash is set at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, which is lower than 3.5 Flash. This reduction in price, combined with improved token efficiency, lowers the overall cost per agentic task, making it more economical to build and run agents.
Performance benchmarks show significant gains over 3.5 Flash. On DeepSWE, 3.6 Flash achieved 49% precision compared to 37% for 3.5 Flash, meaning fewer unwanted code edits and reduced execution loops. On MLE Bench, which measures machine learning research capabilities, 3.6 Flash scored 63.9% versus 49.7% for 3.5 Flash. Computer use capabilities improved on OSWorld-Verified, with 3.6 Flash scoring 83.0% compared to 78.4%. Computer use is now a built-in client side tool available via the Gemini API and Gemini Enterprise. In knowledge work, measured by GDPval-AA v2, 3.6 Flash scored 1421 versus 1349 for 3.5 Flash.
Customers such as Hebbia and Harvey have found 3.6 Flash particularly capable at multimodal tasks, including document parsing, chart and data analysis, and report drafting. The model also demonstrates improved performance in financial data analysis, code migrations, 3D workflow development, and interactive theme studio creation.
Safety has been a focus for 3.6 Flash. The model ships with enhanced Frontier Safety safeguards in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) misuse, as well as cyber offense misuses. These safeguards make the model substantially more resistant to jailbreaks. At the same time, the model has been trained to minimize refusals for beneficial uses. More details are available in the 3.6 Flash model card.
Gemini 3.5 Flash-Lite: Fast and Cost-Effective for High Throughput
Gemini 3.5 Flash-Lite is designed for low-latency tasks and high-throughput workflows such as agentic search and document processing. It is the fastest model in the 3.5 series, running at 350 output tokens per second as measured by Artificial Analysis. Pricing is set at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens. The model offers significantly better quality than 3.1 Flash-Lite, providing a strong price-to-performance ratio for developers and customers running high-volume production traffic.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
3.5 Flash-Lite enables efficient scaling for agentic systems. Across different thinking levels, the model significantly outperforms 3.1 Flash-Lite. Developers can configure the model to prioritize low-latency, low-cost execution for high-volume tasks using minimal and low thinking levels, or engage higher thinking levels to process multi-step subagent workloads. The model also includes computer use as a built-in tool to support agentic tasks across surfaces.
Benchmark results show a significant step up in coding and agentic tasks. On Terminal-Bench 2.1, 3.5 Flash-Lite scored 54% compared to 31% for 3.1 Flash-Lite. On long context tasks measured by GDM-MRCR v2, it scored 72.2% versus 60.1%. On real-world task execution measured by GDPval-AA v2, it scored 1140 versus 642. In many agentic and coding evaluations, 3.5 Flash-Lite even outperforms 3 Flash. On SWE-Bench Pro, it scored 54.2% compared to 49.6% for 3 Flash. On OSWorld-Verified, it scored 74.0% versus 65.1%. This makes 3.5 Flash-Lite a faster and more capable option for workloads that previously relied on 2.5 Flash or 3 Flash.
Early customers have highlighted the model's unique combination of speed, intelligence, and cost efficiency for scaling agentic workflows and data processing tasks. More information is available in the 3.5 Flash-Lite model card.
Gemini 3.5 Flash Cyber: Specialized for Cybersecurity
Gemini 3.5 Flash Cyber is a specialized model built on top of 3.5 Flash and fine-tuned for finding and fixing cybersecurity vulnerabilities. The model is designed to be more efficient than larger models, offering a lower price per token. It operates within CodeMender, a code security agent that uses multiple 3.5 Flash Cyber agents working together to produce a single combined report. On the CyberGym benchmark, 3.5 Flash Cyber reaches competitive performance at the frontier.
Given the dual-use nature of this technology, Google DeepMind has taken an intentional approach to its deployment. The model will be exclusively available to governments and trusted partners via CodeMender as part of a limited-access pilot program. This approach is intended to give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse.
Availability and Next Steps
Gemini 3.6 Flash and 3.5 Flash-Lite are available starting today. Developers can access them in the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity. Enterprises can access the models through the Gemini Enterprise Agent Platform, and 3.6 Flash is also available in the Gemini Enterprise app. For general users, the models are available via the Gemini app. 3.5 Flash-Lite is also rolling out in Google Search.
Google DeepMind is welcoming feedback from developers and customers as they build with 3.6 Flash and 3.5 Flash-Lite. The company also noted that Gemini 3.5 Pro is currently being tested with partners and will be made broadly available when it is ready. Work has already begun on the next generation, with the team starting its most ambitious pre-training run yet for Gemini 4.

