GLM 4.5V
by Z.aitext+image->text2 endpoints
Neura Intelligence Index
46.4/ 100
Rank
#194—
Confidence
mediumBenchmarks
7(18% coverage)
Overview
GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...
Capabilities
Text generation
Image understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Top-P sampling
Stop sequences
Deterministic seed
Modalities
Input
TextImage
Output
Text
Technical Specifications
Context Window
65.5K tokens
Max Output
16.4K tokens
Knowledge Cutoff
2024-12-31
Tokenizer
Other
Uptime (24h)
9835.2%
Parameters
355000000000B
License
mitOpen Weights
Country
🇨🇳CN
Release Date
July 28, 20251y
Supported Parameters
frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningrepetition_penaltyresponse_formatseedstoptemperaturetool_choicetoolstop_ktop_p
Benchmark Performance
coding
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Verified | 64.2 | verified |
| SciCode | 41.7 | verified |
| TerminalBench | 37.5 | verified |
reasoning
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 79.1 | verified |
| Humanity's Last Exam | 14.4 | verified |
agent
| Benchmark | Score | Source |
|---|---|---|
| TAU-Bench Retail | 79.7 | verified |
| BrowseComp | 26.4 | verified |
Performance Over Time
Pricing
Per 1M tokens
Input$0.60
Output$1.80
Cache read$0.11
Added
August 11, 2025
Last synced 9/18/2026