GLM 5.3 FlashX
by Z.aitext+image+video->text1 endpoint
Neura Intelligence Index
94.8/ 100
Rank
#7—
Confidence
lowBenchmarks
2(5% coverage)
Overview
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture (320B total parameters, 18B active), it is suited for efficient coding, visual understanding, and long-horizon agent tasks with a 1M-token context window.
Capabilities
Text generation
Image understanding
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling
Modalities
Input
TextImageVideo
Output
Text
Technical Specifications
Context Window
1.0M tokens
Max Output
131.1K tokens
Tokenizer
Other
Uptime (24h)
9999.9%
Parameters
753000000000B
License
glm_5_3Open Weights
Country
🇨🇳CN
Release Date
August 14, 20261mo
Supported Parameters
include_reasoningmax_tokensreasoningreasoning_effortresponse_formattemperaturetool_choicetoolstop_ktop_p
Benchmark Performance
reasoning
| Benchmark | Score | Source |
|---|---|---|
| Humanity's Last Exam | 62.5 | verified |
agent
| Benchmark | Score | Source |
|---|---|---|
| Toolathlon | 73 | verified |
Performance Over Time
Pricing
Per 1M tokens
Input$0.37
Output$1.25
Cache read$0.07
Added
September 18, 2026
Last synced 9/21/2026