Models/NVIDIA/Nemotron 3 Ultra

Nemotron 3 Ultra

NVIDIA · Open weights · Released Jun 4, 2026 · nvidia/nemotron-3-ultra-550b-a55b
Best $0.50 in / $2.20 out via DeepInfra · fp4
CompareEstimate cost
262K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 4 providersUSD / MTOK · PRICES 23 MINUTES AGO · SPEED 3 DAYS AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
PROVIDERINPUTOUTPUTCACHECONTEXTQUANTTTFTTPSUPTIME
TTFT/TPS: median (P50) over trailing 504h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · ✕ = probed, endpoint returned no measurable stream · methodology v1.1 · measured 3 days ago
SERVE THIS MODEL AND NOT LISTED? LIST YOUR INFERENCE →
Pricing variants · USD / MTok in / out
STANDARD
$0.50 / $2.20
per MTok in / out
FREE
Free / Free
per MTok in / out
CACHE READ $0.10
Recent changesRSS ↗
Aug 29, 2026endpoint removed · together4 → 3
Aug 29, 2026context changed512288 → 262144
About

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Specifications
Context window262,144
Max output16,384
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedJun 4, 2026
TokenizerOther
Open weights
Downloads623.6K
Likes361
Licenseother
safetensorspytorch
Embed badge
Nemotron 3 Ultra price and speed badge
[![Nemotron 3 Ultra](https://modelindex.ai/badge/nvidia/nemotron-3-ultra-550b-a55b)](https://modelindex.ai/models/nvidia/nemotron-3-ultra-550b-a55b)
Sources
OpenRouter API23 minutes ago
AGGREGATOR
Hugging Face23 minutes ago
COMMUNITY
ModelIndex benchmarks14 days ago
MEASURED
Verified 23 minutes ago · list prices, not negotiated rates.