Inception

Mercury 2.5

by Inceptiontext->text1 endpoint
Neura Intelligence Index
52.2/ 100
Rank
#172
Confidence
low
Benchmarks
3(8% coverage)

Overview

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents. Read more in the blog post.

Capabilities

Text generation
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Stop sequences

Modalities

Input
Text
Output
Text

Technical Specifications

Context Window
260K tokens
Max Output
65.5K tokens
Tokenizer
Other
Uptime (24h)
9999.7%
Country
🇺🇸US
Release Date
February 24, 20266mo

Supported Parameters

include_reasoningmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetools

Benchmark Performance

coding

BenchmarkScoreSource
SciCode
38
verified

math

BenchmarkScoreSource
AIME 2025
91.1
verified

reasoning

BenchmarkScoreSource
GPQA Diamond
74
verified

Performance Over Time

Pricing

Per 1M tokens
Input$0.04
Output$0.15
Cache read$0.0040
Added
September 8, 2026
Last synced 9/18/2026