Qwen3 VL 30B A3B Instruct
by Qwentext+image->text4 endpoints
Neura Intelligence Index
37.3/ 100
Rank
#239—
Confidence
mediumBenchmarks
7(18% coverage)
Overview
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...
Capabilities
Text generation
Image understanding
Tool use / Function calling
Structured output
Long context
Top-P sampling
Stop sequences
Deterministic seed
Log probabilities
Modalities
Input
TextImage
Output
Text
Technical Specifications
Context Window
262.1K tokens
Max Output
32.8K tokens
Knowledge Cutoff
2025-03-31
Tokenizer
Qwen3
Uptime (24h)
9995.4%
Parameters
31000000000B
License
apache_2_0Open Weights
Country
🇨🇳CN
Release Date
September 22, 202512mo
Supported Parameters
frequency_penaltylogit_biaslogprobsmax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Benchmark Performance
general
| Benchmark | Score | Source |
|---|---|---|
| SimpleQA | 27 | verified |
math
| Benchmark | Score | Source |
|---|---|---|
| AIME 2025 | 69.3 | verified |
reasoning
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 70.4 | verified |
multimodal
| Benchmark | Score | Source |
|---|---|---|
| ScreenSpot Pro | 60.5 | verified |
| MMMU-Pro | 60.4 | verified |
| CharXiv Reasoning | 48.9 | verified |
agent
| Benchmark | Score | Source |
|---|---|---|
| OSWorld | 30.3 | verified |
Performance Over Time
Pricing
Per 1M tokens
Input$0.13
Output$0.52
Added
October 6, 2025
Last synced 9/16/2026