Gemini 2.5 Flash Lite (batch)
by Googletext+image+file+audio+video->text2 endpoints
Overview
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, "thinking" (i.e. multi-pass reasoning) is disabled to prioritize speed, but developers can enable it via the Reasoning API parameter to selectively trade off cost for intelligence.
Capabilities
Text generation
Image understanding
Audio processing
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
File input
Top-P sampling
Stop sequences
Deterministic seed
Modalities
Input
TextImageFileAudioVideo
Output
Text
Technical Specifications
Context Window
1.0M tokens
Max Output
65.5K tokens
Knowledge Cutoff
2025-01-31
Tokenizer
Gemini
Supported Parameters
include_reasoningmax_tokensreasoningresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
Pricing
Per 1M tokens
Input$0.05
Output$0.20
Cache read$0.01
Audio$0.15
Reasoning$0.20
Image$0.05
Added
July 22, 2025
Last synced 9/19/2026