Qwen

Qwen3 VL 235B A22B Instruct

by Qwentext+image->text5 endpoints
Neura Intelligence Index
54.7/ 100
Rank
#164
Confidence
medium
Benchmarks
6(15% coverage)

Overview

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table...

Capabilities

Text generation
Image understanding
Tool use / Function calling
Structured output
Prompt caching
Long context
Top-P sampling
Stop sequences
Deterministic seed
Log probabilities

Modalities

Input
TextImage
Output
Text

Technical Specifications

Context Window
262.1K tokens
Max Output
32.8K tokens
Knowledge Cutoff
2025-03-31
Tokenizer
Qwen3
Uptime (24h)
9970.6%
Parameters
236000000000B
License
apache_2_0Open Weights
Country
🇨🇳CN
Release Date
September 22, 202512mo

Supported Parameters

frequency_penaltylogit_biaslogprobsmax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Benchmark Performance

general

BenchmarkScoreSource
SimpleQA
51.9
verified

math

BenchmarkScoreSource
AIME 2025
74.7
verified

multimodal

BenchmarkScoreSource
MMMU-Pro
68.1
verified
CharXiv Reasoning
62.1
verified
ScreenSpot Pro
62
verified

agent

BenchmarkScoreSource
OSWorld
66.7
verified

Performance Over Time

Pricing

Per 1M tokens
Input$0.21
Output$1.90
Cache read$0.10
Added
September 23, 2025
Last synced 9/18/2026