Gemini 2.5 Flash Lite
by Googletext+image+file+audio+video->text5 endpoints
Neura Intelligence Index
18.8/ 100
Rank
#296—
Confidence
mediumBenchmarks
6(15% coverage)
Overview
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
Capabilities
Text generation
Image understanding
Audio processing
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
File input
Top-P sampling
Stop sequences
Deterministic seed
Modalities
Input
TextImageFileAudioVideo
Output
Text
Technical Specifications
Context Window
1.0M tokens
Max Output
65.5K tokens
Knowledge Cutoff
2025-01-31
Tokenizer
Gemini
Uptime (24h)
10000.0%
License
creative_commons_attribution_4_0_licenseOpen Weights
Country
🇺🇸US
Release Date
June 17, 20251y
Supported Parameters
include_reasoningmax_tokensreasoningresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
Benchmark Performance
general
| Benchmark | Score | Source |
|---|---|---|
| SimpleQA | 10.7 | verified |
coding
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Verified | 31.6 | verified |
math
| Benchmark | Score | Source |
|---|---|---|
| AIME 2025 | 49.8 | verified |
reasoning
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 64.6 | verified |
| Humanity's Last Exam | 5.1 | verified |
multimodal
| Benchmark | Score | Source |
|---|---|---|
| MMMU | 72.9 | verified |
Performance Over Time
Pricing
Per 1M tokens
Input$0.10
Output$0.40
Cache read$0.01
Cache write$0.08
Audio$0.30
Reasoning$0.40
Image$0.10
Added
July 22, 2025
Last synced 9/18/2026