MiMo-V2.5
by Xiaomitext+image+audio+video->text5 endpoints
Neura Intelligence Index
54.3/ 100
Rank
#166—
Confidence
lowBenchmarks
3(8% coverage)
Overview
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
Capabilities
Text generation
Image understanding
Audio processing
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling
Stop sequences
Deterministic seed
Modalities
Input
TextAudioImageVideo
Output
Text
Technical Specifications
Context Window
1.1M tokens
Max Output
131.1K tokens
Tokenizer
Other
Uptime (24h)
9895.1%
Parameters
310775040000B
License
mitOpen Weights
Country
🇨🇳CN
Release Date
April 22, 20265mo
Supported Parameters
frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Benchmark Performance
coding
| Benchmark | Score | Source |
|---|---|---|
| SWE-Bench Pro | 56.1 | verified |
multimodal
| Benchmark | Score | Source |
|---|---|---|
| CharXiv Reasoning | 81 | verified |
| MMMU-Pro | 77.9 | verified |
Performance Over Time
Pricing
Per 1M tokens
Input$0.14
Output$0.28
Cache read$0.0028
Added
April 22, 2026
Last synced 9/19/2026