Inclusion AI

inclusionAI: Ling 3.0 Flash

by Inclusion AItext->text2 endpoints
Neura Intelligence Index
47.8/ 100
Rank
#186
Confidence
low
Benchmarks
3(8% coverage)

Overview

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token.

The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.

Capabilities

Text generation
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling
Stop sequences
Deterministic seed
Log probabilities

Modalities

Input
Text
Output
Text

Technical Specifications

Context Window
262.1K tokens
Max Output
32.8K tokens
Tokenizer
Other
Uptime (24h)
9998.7%
Parameters
124000000000B
License
unknownOpen Weights
Release Date
August 4, 20261mo

Supported Parameters

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstoptemperaturetool_choicetoolstop_ktop_logprobstop_p

Benchmark Performance

coding

BenchmarkScoreSource
SWE-Bench Pro
56.6
verified

agent

BenchmarkScoreSource
BrowseComp
72.2
verified
MCP Atlas
65.5
verified

Performance Over Time

Pricing

Per 1M tokens
Input$0.02
Output$0.06
Cache read$0.0042
Added
July 23, 2026
Last synced 9/20/2026