Z.ai

GLM 5.3 FlashX

by Z.aitext+image+video->text1 endpoint
Neura Intelligence Index
94.8/ 100
Rank
#7
Confidence
low
Benchmarks
2(5% coverage)

Overview

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture (320B total parameters, 18B active), it is suited for efficient coding, visual understanding, and long-horizon agent tasks with a 1M-token context window.

Capabilities

Text generation
Image understanding
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling

Modalities

Input
TextImageVideo
Output
Text

Technical Specifications

Context Window
1.0M tokens
Max Output
131.1K tokens
Tokenizer
Other
Uptime (24h)
9999.9%
Parameters
753000000000B
License
glm_5_3Open Weights
Country
🇨🇳CN
Release Date
August 14, 20261mo

Supported Parameters

include_reasoningmax_tokensreasoningreasoning_effortresponse_formattemperaturetool_choicetoolstop_ktop_p

Benchmark Performance

reasoning

BenchmarkScoreSource
Humanity's Last Exam
62.5
verified

agent

BenchmarkScoreSource
Toolathlon
73
verified

Performance Over Time

Pricing

Per 1M tokens
Input$0.37
Output$1.25
Cache read$0.07
Added
September 18, 2026
Last synced 9/21/2026