Qwen3 VL 32B Instruct
by Qwentext+image->text1 endpoint
Neura Intelligence Index
37.6/ 100
Rank
#236—
Confidence
mediumBenchmarks
6(15% coverage)
Overview
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...
Capabilities
Text generation
Image understanding
Tool use / Function calling
Structured output
Long context
Top-P sampling
Stop sequences
Deterministic seed
Log probabilities
Modalities
Input
TextImage
Output
Text
Technical Specifications
Context Window
131.1K tokens
Max Output
32.8K tokens
Tokenizer
Qwen
Uptime (24h)
9999.8%
Parameters
33000000000B
License
apache_2_0Open Weights
Country
🇨🇳CN
Release Date
September 22, 202512mo
Supported Parameters
frequency_penaltylogprobsmax_tokenspresence_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Benchmark Performance
math
| Benchmark | Score | Source |
|---|---|---|
| AIME 2025 | 66.2 | verified |
reasoning
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 68.9 | verified |
multimodal
| Benchmark | Score | Source |
|---|---|---|
| MMMU-Pro | 65.3 | verified |
| CharXiv Reasoning | 62.8 | verified |
| ScreenSpot Pro | 57.9 | verified |
agent
| Benchmark | Score | Source |
|---|---|---|
| OSWorld | 32.6 | verified |
Performance Over Time
Pricing
Per 1M tokens
Input$0.10
Output$0.42
Added
October 23, 2025
Last synced 9/20/2026