Xiaomi

MiMo-V2.5

by Xiaomitext+image+audio+video->text5 endpoints
Neura Intelligence Index
54.3/ 100
Rank
#166
Confidence
low
Benchmarks
3(8% coverage)

Overview

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

Capabilities

Text generation
Image understanding
Audio processing
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling
Stop sequences
Deterministic seed

Modalities

Input
TextAudioImageVideo
Output
Text

Technical Specifications

Context Window
1.1M tokens
Max Output
131.1K tokens
Tokenizer
Other
Uptime (24h)
9895.1%
Parameters
310775040000B
License
mitOpen Weights
Country
🇨🇳CN
Release Date
April 22, 20265mo

Supported Parameters

frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p

Benchmark Performance

coding

BenchmarkScoreSource
SWE-Bench Pro
56.1
verified

multimodal

BenchmarkScoreSource
CharXiv Reasoning
81
verified
MMMU-Pro
77.9
verified

Performance Over Time

Pricing

Per 1M tokens
Input$0.14
Output$0.28
Cache read$0.0028
Added
April 22, 2026
Last synced 9/19/2026