Ling 3.0 Flash
InclusionAI · Open weights · Released Jul 23, 2026 · inclusionai/ling-3.0-flash
Best $0.021 in / $0.063 out via Novita
262K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 2 providersUSD / MTOK · PRICES 23 MINUTES AGO · SPEED 5 DAYS AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.021 | $0.063 | $0.004 | 262K | — | 0.98s | 372 tok/s | 100.0% | |
| $0.06 | $0.18 | $0.012 | 131K | bf16 | ✕ | ✕ | 97.3% |
TTFT/TPS: median (P50) over trailing 504h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · ✕ = probed, endpoint returned no measurable stream · methodology v1.1 · measured 5 days ago
SERVE THIS MODEL AND NOT LISTED? LIST YOUR INFERENCE →
Pricing variants · USD / MTok in / out
STANDARD
$0.021 / $0.063
per MTok in / out
CACHE READ $0.004
About
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Specifications
Context window262,144
Max output32,768
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedJul 23, 2026
TokenizerOther
WeightsHugging Face ↗
Embed badge
[](https://modelindex.ai/models/inclusionai/ling-3.0-flash)Sources
OpenRouter API23 minutes ago
AGGREGATORHugging Face23 minutes ago
COMMUNITYModelIndex benchmarks10 days ago
MEASUREDVerified 23 minutes ago · list prices, not negotiated rates.