Nemotron 3 Ultra
NVIDIA · Open weights · Released Jun 4, 2026 · nvidia/nemotron-3-ultra-550b-a55b
512K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 4 providersUSD / MTOK · VERIFIED 4 MINUTES AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.50 | $2.20 | $0.10 | 262K | fp4 | 3.4s | — | 100.0% | |
| $0.60 | $2.40 | $0.12 | 203K | fp4 | 0.58s | 234 tok/s | 99.7% | |
| $0.60 | $3.60 | $0.20 | 512K | — | 0.64s | 153 tok/s | 98.3% | |
| $0.625 | $3.13 | $0.188 | 256K | fp8 | 3.6s | — | 98.1% |
TTFT/TPS: median (P50) over trailing 72h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · methodology v1.1 · measured 44 hours ago
Pricing variants · USD / MTok in / out
STANDARD
$0.50 / $2.20
per MTok in / out
BATCH
$0.60 / $3.60
async batch
FREE
Free / Free
per MTok in / out
CACHE READ $0.10
About
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Specifications
Context window512,288
Max output—
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedJun 4, 2026
TokenizerOther
WeightsHugging Face ↗
Open weights
safetensorspytorch
Related
Embed badge
[](https://modelindex.ai/models/nvidia/nemotron-3-ultra-550b-a55b)Sources
OpenRouter API4 minutes ago
AGGREGATORHugging Face4 minutes ago
COMMUNITYModelIndex benchmarks44 hours ago
MEASUREDVerified 4 minutes ago · list prices, not negotiated rates.