Ling-3.0-flash
InclusionAI · Open weights · Released Jul 23, 2026 · inclusionai/ling-3.0-flash
262K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 3 providersUSD / MTOK · VERIFIED 4 MINUTES AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.021 | $0.063 | $0.004 | 262K | — | 1s | 388 tok/s | 100.0% | |
| $0.042 | $0.126 | $0.008 | 262K | — | — | — | 99.9% | |
| $0.06 | $0.18 | $0.012 | 131K | bf16 | 0.84s | 83 tok/s | 100.0% |
TTFT/TPS: median (P50) over trailing 72h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · methodology v1.1 · measured 44 hours ago
Pricing variants · USD / MTok in / out
STANDARD
$0.021 / $0.063
per MTok in / out
CACHE READ $0.004
About
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Specifications
Context window262,144
Max output32,768
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedJul 23, 2026
TokenizerOther
WeightsHugging Face ↗
Embed badge
[](https://modelindex.ai/models/inclusionai/ling-3.0-flash)Sources
OpenRouter API4 minutes ago
AGGREGATORHugging Face4 minutes ago
COMMUNITYModelIndex benchmarks44 hours ago
MEASUREDVerified 4 minutes ago · list prices, not negotiated rates.