Google

Gemini 2.5 Flash Lite (batch)

by Googletext+image+file+audio+video->text2 endpoints

Overview

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, "thinking" (i.e. multi-pass reasoning) is disabled to prioritize speed, but developers can enable it via the Reasoning API parameter to selectively trade off cost for intelligence.

Capabilities

Text generation
Image understanding
Audio processing
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
File input
Top-P sampling
Stop sequences
Deterministic seed

Modalities

Input
TextImageFileAudioVideo
Output
Text

Technical Specifications

Context Window
1.0M tokens
Max Output
65.5K tokens
Knowledge Cutoff
2025-01-31
Tokenizer
Gemini

Supported Parameters

include_reasoningmax_tokensreasoningresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p

Pricing

Per 1M tokens
Input$0.05
Output$0.20
Cache read$0.01
Audio$0.15
Reasoning$0.20
Image$0.05
Added
July 22, 2025
Last synced 9/19/2026