Nemotron 3.5 Lightning
NEWNVIDIA · Open weights · Released Aug 11, 2026 · nvidia/nemotron-3.5-lightning
262K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 2 providersUSD / MTOK · VERIFIED 4 MINUTES AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.08 | $0.20 | $0.04 | 262K | bf16 | 0.29s | 103 tok/s | 100.0% | |
| CWCoreWeave | $0.10 | $0.25 | $0.05 | 262K | bf16 | 0.53s | 397 tok/s | 100.0% |
TTFT/TPS: median (P50) over trailing 72h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · methodology v1.1 · measured 44 hours ago
Pricing variants · USD / MTok in / out
STANDARD
$0.08 / $0.20
per MTok in / out
FREE
Free / Free
per MTok in / out
CACHE READ $0.04
About
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
Specifications
Context window262,144
Max output131,072
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedAug 11, 2026
TokenizerOther
WeightsHugging Face ↗
Open weights
safetensorspytorch
Related
Embed badge
[](https://modelindex.ai/models/nvidia/nemotron-3.5-lightning)Sources
OpenRouter API4 minutes ago
AGGREGATORHugging Face4 minutes ago
COMMUNITYModelIndex benchmarks44 hours ago
MEASUREDVerified 4 minutes ago · list prices, not negotiated rates.