Qwen3 VL 8B Instruct
by Qwentext+image->text2 endpoints
Neura Intelligence Index
12.3/ 100
Rank
#317—
Confidence
mediumBenchmarks
5(13% coverage)
Overview
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
Capabilities
Text generation
Image understanding
Tool use / Function calling
Structured output
Long context
Top-P sampling
Stop sequences
Deterministic seed
Log probabilities
Modalities
Input
ImageText
Output
Text
Technical Specifications
Context Window
262.1K tokens
Max Output
32.8K tokens
Tokenizer
Qwen3
Uptime (24h)
9998.4%
Parameters
9000000000B
License
apache_2_0Open Weights
Country
🇨🇳CN
Release Date
September 22, 202512mo
Supported Parameters
frequency_penaltylogit_biaslogprobsmax_tokenspresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Benchmark Performance
math
| Benchmark | Score | Source |
|---|---|---|
| AIME 2025 | 45.9 | verified |
multimodal
| Benchmark | Score | Source |
|---|---|---|
| MMMU-Pro | 55.9 | verified |
| ScreenSpot Pro | 54.6 | verified |
| CharXiv Reasoning | 46.4 | verified |
agent
| Benchmark | Score | Source |
|---|---|---|
| OSWorld | 33.9 | verified |
Performance Over Time
Pricing
Per 1M tokens
Input$0.12
Output$0.46
Added
October 14, 2025
Last synced 9/16/2026