Mercury 2.5
by Inceptiontext->text1 endpoint
Neura Intelligence Index
52.2/ 100
Rank
#172—
Confidence
lowBenchmarks
3(8% coverage)
Overview
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents. Read more in the blog post.
Capabilities
Text generation
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Stop sequences
Modalities
Input
Text
Output
Text
Technical Specifications
Context Window
260K tokens
Max Output
65.5K tokens
Tokenizer
Other
Uptime (24h)
9999.7%
Country
🇺🇸US
Release Date
February 24, 20266mo
Supported Parameters
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetools
Benchmark Performance
coding
| Benchmark | Score | Source |
|---|---|---|
| SciCode | 38 | verified |
math
| Benchmark | Score | Source |
|---|---|---|
| AIME 2025 | 91.1 | verified |
reasoning
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 74 | verified |
Performance Over Time
Pricing
Per 1M tokens
Input$0.04
Output$0.15
Cache read$0.0040
Added
September 8, 2026
Last synced 9/18/2026