Qwen3.5-9B (batch)
by Qwentext+image+video->text1 endpoint
Overview
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design with early fusion of multimodal tokens, allowing the model to process and reason across text and images within the same context.
Capabilities
Text generation
Image understanding
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Long context
Top-P sampling
Stop sequences
Modalities
Input
TextImageVideo
Output
Text
Technical Specifications
Context Window
262.1K tokens
Max Output
235.9K tokens
Tokenizer
Qwen3
Supported Parameters
frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Pricing
Per 1M tokens
Input$0.17
Output$0.25
Added
March 10, 2026
Last synced 9/19/2026