DeepSeek V4 Flash Vision Exp (batch)
by DeepSeektext+image->text1 endpoint
Overview
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge. It is a sparse mixture-of-experts model with 13B active parameters out of 284B total.
It is suited for document and chart understanding, visual question answering, and multimodal agent workflows that interleave text and images.
Capabilities
Text generation
Image understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling
Stop sequences
Log probabilities
Modalities
Input
TextImage
Output
Text
Technical Specifications
Context Window
1.0M tokens
Max Output
943.7K tokens
Tokenizer
DeepSeek
Supported Parameters
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokenspresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Pricing
Per 1M tokens
Input$0.11
Output$0.33
Cache read$0.0035
Added
August 21, 2026
Last synced 9/16/2026