inclusionAI: Ling 3.0 Flash
by Inclusion AItext->text2 endpoints
Neura Intelligence Index
47.8/ 100
Rank
#186—
Confidence
lowBenchmarks
3(8% coverage)
Overview
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token.
The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
Capabilities
Text generation
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling
Stop sequences
Deterministic seed
Log probabilities
Modalities
Input
Text
Output
Text
Technical Specifications
Context Window
262.1K tokens
Max Output
32.8K tokens
Tokenizer
Other
Uptime (24h)
9998.7%
Parameters
124000000000B
License
unknownOpen Weights
Release Date
August 4, 20261mo
Supported Parameters
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstoptemperaturetool_choicetoolstop_ktop_logprobstop_p
Benchmark Performance
coding
| Benchmark | Score | Source |
|---|---|---|
| SWE-Bench Pro | 56.6 | verified |
agent
| Benchmark | Score | Source |
|---|---|---|
| BrowseComp | 72.2 | verified |
| MCP Atlas | 65.5 | verified |
Performance Over Time
Pricing
Per 1M tokens
Input$0.02
Output$0.06
Cache read$0.0042
Added
July 23, 2026
Last synced 9/20/2026