Nemotron 3 Ultra
NVIDIA · Open weights · Released Jun 4, 2026 · nvidia/nemotron-3-ultra-550b-a55b
Best $0.50 in / $2.20 out via DeepInfra · fp4
262K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 4 providersUSD / MTOK · PRICES 23 MINUTES AGO · SPEED 3 DAYS AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.50 | $2.20 | $0.10 | 262K | fp4 | 1s | 267 tok/s | 99.9% | |
| $0.60 | $2.40 | $0.12 | 203K | fp4 | 0.51s | 341 tok/s | 100.0% | |
| $0.60 | $2.40 | $0.12 | 203K | fp4 | — | — | 100.0% | |
| $0.625 | $3.13 | $0.188 | 256K | fp8 | ✕ | ✕ | 91.1% |
TTFT/TPS: median (P50) over trailing 504h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · ✕ = probed, endpoint returned no measurable stream · methodology v1.1 · measured 3 days ago
SERVE THIS MODEL AND NOT LISTED? LIST YOUR INFERENCE →
Pricing variants · USD / MTok in / out
STANDARD
$0.50 / $2.20
per MTok in / out
FREE
Free / Free
per MTok in / out
CACHE READ $0.10
Recent changesRSS ↗
Aug 29, 2026endpoint removed · together4 → 3
Aug 29, 2026context changed512288 → 262144
About
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Specifications
Context window262,144
Max output16,384
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedJun 4, 2026
TokenizerOther
WeightsHugging Face ↗
Open weights
safetensorspytorch
Related
Embed badge
[](https://modelindex.ai/models/nvidia/nemotron-3-ultra-550b-a55b)Sources
OpenRouter API23 minutes ago
AGGREGATORHugging Face23 minutes ago
COMMUNITYModelIndex benchmarks14 days ago
MEASUREDVerified 23 minutes ago · list prices, not negotiated rates.