Z.ai

GLM 5V Turbo

by Z.aitext+image+video->text1 endpoint
Neura Intelligence Index
76.5/ 100
Rank
#48
Confidence
low
Benchmarks
1(3% coverage)

Overview

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

Capabilities

Text generation
Image understanding
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling

Modalities

Input
ImageTextVideo
Output
Text

Technical Specifications

Context Window
202.8K tokens
Max Output
131.1K tokens
Tokenizer
Other
Uptime (24h)
10000.0%
Country
🇨🇳CN
Release Date
April 2, 20265mo

Supported Parameters

include_reasoningmax_tokensreasoningresponse_formattemperaturetool_choicetoolstop_ktop_p

Benchmark Performance

agent

BenchmarkScoreSource
OSWorld
62.3
verified

Performance Over Time

Pricing

Per 1M tokens
Input$1.20
Output$4.00
Cache read$0.24
Added
April 1, 2026
Last synced 9/16/2026