Qwen

Qwen3 VL 30B A3B Instruct

by Qwentext+image->text4 endpoints
Neura Intelligence Index
37.3/ 100
Rank
#239
Confidence
medium
Benchmarks
7(18% coverage)

Overview

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...

Capabilities

Text generation
Image understanding
Tool use / Function calling
Structured output
Long context
Top-P sampling
Stop sequences
Deterministic seed
Log probabilities

Modalities

Input
TextImage
Output
Text

Technical Specifications

Context Window
262.1K tokens
Max Output
32.8K tokens
Knowledge Cutoff
2025-03-31
Tokenizer
Qwen3
Uptime (24h)
9995.4%
Parameters
31000000000B
License
apache_2_0Open Weights
Country
🇨🇳CN
Release Date
September 22, 202512mo

Supported Parameters

frequency_penaltylogit_biaslogprobsmax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Benchmark Performance

general

BenchmarkScoreSource
SimpleQA
27
verified

math

BenchmarkScoreSource
AIME 2025
69.3
verified

reasoning

BenchmarkScoreSource
GPQA Diamond
70.4
verified

multimodal

BenchmarkScoreSource
ScreenSpot Pro
60.5
verified
MMMU-Pro
60.4
verified
CharXiv Reasoning
48.9
verified

agent

BenchmarkScoreSource
OSWorld
30.3
verified

Performance Over Time

Pricing

Per 1M tokens
Input$0.13
Output$0.52
Added
October 6, 2025
Last synced 9/16/2026