Z.ai

GLM 4.6V

by Z.aitext+image+video->text2 endpoints
Neura Intelligence Index
57.3/ 100
Rank
#152
Confidence
medium
Benchmarks
6(15% coverage)

Overview

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

Capabilities

Text generation
Image understanding
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling
Stop sequences
Deterministic seed

Modalities

Input
ImageTextVideo
Output
Text

Technical Specifications

Context Window
131.1K tokens
Max Output
32.8K tokens
Tokenizer
Other
Uptime (24h)
9868.9%
Parameters
357000000000B
License
mitOpen Weights
Country
🇨🇳CN
Release Date
September 30, 202511mo

Supported Parameters

frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningrepetition_penaltyresponse_formatseedstoptemperaturetool_choicetoolstop_ktop_p

Benchmark Performance

coding

BenchmarkScoreSource
SWE-bench Verified
68
verified
TerminalBench
40.5
verified

math

BenchmarkScoreSource
AIME 2025
93.9
verified

reasoning

BenchmarkScoreSource
GPQA Diamond
81
verified
Humanity's Last Exam
17.2
verified

agent

BenchmarkScoreSource
BrowseComp
45.1
verified

Performance Over Time

Pricing

Per 1M tokens
Input$0.30
Output$0.90
Cache read$0.06
Added
December 8, 2025
Last synced 9/19/2026