AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Neura News

Neura News

Neura Market Editorial

July 21, 20265 min read

Originally reported by blog.google

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has introduced three new models in its Gemini Flash series, aiming to give developers and enterprises more efficient tools for building AI agents at scale. The releases include Gemini 3.6 Flash, 3.5 Flash-Lite, and a specialized cybersecurity model called 3.5 Flash Cyber.

According to Tulsee Doshi, Senior Director of Product Management at Google, the new models are designed to meet the need for higher token efficiency, lower latency, and more reliable performance in production AI agent workflows. The Flash series has been built to balance efficiency and quality, enabling developers to scale agentic systems.

3.6 Flash: Better Performance with Fewer Tokens

Gemini 3.6 Flash builds directly on feedback from users of 3.5 Flash. It delivers a step up in coding and knowledge work while also improving token efficiency. According to the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash. It also requires fewer reasoning steps and tool calls to complete multi-step workflows.

The model is priced lower than its predecessor. Input tokens cost $1.50 per million, and output tokens cost $7.50 per million. This reduces the overall cost per agentic task, making agents more cost-effective to build and run.

Performance benchmarks show clear gains over 3.5 Flash. On the DeepSWE benchmark from Datacurve, 3.6 Flash achieved 49% precision compared to 37% for 3.5 Flash, with fewer unwanted code edits and reduced execution loops. On MLE Bench, which measures ML research capability, it scored 63.9% versus 49.7%. Computer use capabilities improved on OSWorld-Verified, reaching 83.0% compared to 78.4%. Computer use is now a built-in client side tool available through the Gemini API and Gemini Enterprise.

In knowledge work, 3.6 Flash outperformed 3.5 Flash on the GDPval-AA v2 benchmark, scoring 1421 versus 1349. Customers such as Hebbia and Harvey have found the model particularly capable at multimodal tasks like document parsing, chart and data analysis, and report drafting.

The model also showed gains in financial data analysis, code migrations using multi-agent orchestration, and 3D workflows. One example involved developing a photographic texture extractor for 3D workflows using the Gemini App. Another showed the model building interactive theme studios using the tldraw offline editor.

Google emphasized that 3.6 Flash ships with enhanced Frontier Safety safeguards in the domains of chemical, biological, radiological, and nuclear (CBRN) threats, as well as cyber offense misuses. These safeguards make the model substantially more resistant to jailbreaks while minimizing refusals for beneficial uses. More details are available in the 3.6 Flash model card.

3.5 Flash-Lite: Speed and Scale for Agentic Workflows

Gemini 3.5 Flash-Lite is designed for both low-latency tasks and high-throughput workloads such as agentic search and document processing. It is the fastest model in the 3.5 series, running at 350 output tokens per second according to the Artificial Analysis Index.

Pricing is set at $0.30 per million input tokens and $2.50 per million output tokens. The model offers significantly better quality than 3.1 Flash-Lite, providing a strong price-to-performance ratio for developers and customers running high volume production traffic.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Across thinking levels, 3.5 Flash-Lite significantly outperforms 3.1 Flash-Lite. Developers can configure the model to prioritize low-latency, low-cost execution for high-volume tasks using minimal or low thinking levels, or engage higher thinking levels for multi-step subagent workloads. The model also includes computer use as a built-in tool to support agentic tasks across different surfaces.

Benchmark results show a significant step up in coding and agentic tasks. On Terminal-Bench 2.1, it scored 54% compared to 31% for 3.1 Flash-Lite. On the long context benchmark GDM-MRCR v2, it achieved 72.2% versus 60.1%. On the real-world task execution benchmark GDPval-AA v2, it scored 1140 compared to 642.

Notably, 3.5 Flash-Lite even outperforms the older 3 Flash model on several agentic and coding evaluations. On SWE-Bench Pro, it scored 54.2% versus 49.6% for 3 Flash. On OSWorld-Verified, it achieved 74.0% compared to 65.1%. This makes it a faster and more capable option for workloads that previously ran on 2.5 Flash or 3 Flash.

Early customers have highlighted the model's unique combination of speed, intelligence, and cost efficiency for scaling agentic workflows and data processing tasks. More information is available in the 3.5 Flash-Lite model card.

3.5 Flash Cyber: Specialized for Security

Google also introduced Gemini 3.5 Flash Cyber, a specialized model built on top of 3.5 Flash and fine-tuned for finding and fixing cybersecurity vulnerabilities. The model is designed to detect, validate, and patch code security issues at scale, at a lower price per token than larger models.

Within CodeMender, Google's code security agent, multiple 3.5 Flash Cyber agents work together to produce a single combined report. On the popular CyberGym benchmark, the model reaches competitive performance at the frontier.

Given the dual-use nature of this technology, Google has taken a cautious approach to deployment. The model will be exclusively available to governments and trusted partners via CodeMender as part of a limited-access pilot program. This is intended to give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse.

Availability and Next Steps

Gemini 3.6 Flash and 3.5 Flash-Lite are available starting today. Developers can access them through the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity. Enterprises can use the models through the Gemini Enterprise Agent Platform, and 3.6 Flash is also available in the Gemini Enterprise app. For consumers, both models are available via the Gemini app. 3.5 Flash-Lite is also rolling out in Google Search.

Google noted that Gemini 3.5 Pro is currently being tested with partners and will be made broadly available as soon as it is ready. The company also revealed that it has started its most ambitious pre-training run yet for Gemini 4 and is excited by the progress.

Related on Neura Market

More from Neura News

AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read