Models/NVIDIA/Nemotron 3 Ultra

Nemotron 3 Ultra

NVIDIA · Open weights · Released Jun 4, 2026 · nvidia/nemotron-3-ultra-550b-a55b
CompareEstimate cost
512K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 4 providersUSD / MTOK · VERIFIED 4 MINUTES AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
PROVIDERINPUTOUTPUTCACHECONTEXTQUANTTTFTTPSUPTIME
TTFT/TPS: median (P50) over trailing 72h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · methodology v1.1 · measured 44 hours ago
Pricing variants · USD / MTok in / out
STANDARD
$0.50 / $2.20
per MTok in / out
BATCH
$0.60 / $3.60
async batch
FREE
Free / Free
per MTok in / out
CACHE READ $0.10
About

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Specifications
Context window512,288
Max output
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedJun 4, 2026
TokenizerOther
Open weights
Downloads402.6K
Likes329
Licenseother
safetensorspytorch
Embed badge
Nemotron 3 Ultra price and speed badge[![Nemotron 3 Ultra](https://modelindex.ai/badge/nvidia/nemotron-3-ultra-550b-a55b)](https://modelindex.ai/models/nvidia/nemotron-3-ultra-550b-a55b)
Sources
OpenRouter API4 minutes ago
AGGREGATOR
Hugging Face4 minutes ago
COMMUNITY
ModelIndex benchmarks44 hours ago
MEASURED
Verified 4 minutes ago · list prices, not negotiated rates.