NVIDIA
NVIDIA Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is a NVIDIA model with 1M context, $0.50 in / $2.20 out per million tokens, last verified September 2, 2026. Open weights: yes. Public-board scores are linked to the publisher; vendor-only figures stay in the watch-out, not on the board.
Nemotron 3 Ultra is NVIDIA's open-weight (550B/55B) frontier MoE with a 1M-token context.
Catalog record checked September 2, 2026Individual provider fields may change
| Specification | NVIDIA Nemotron 3 Ultra |
|---|---|
| Provider | NVIDIA |
| Tier | Frontier |
| Context window | 1M |
| Max output | Not verifiedUnverified |
| Input / 1M tokens | $0.50 |
| Output / 1M tokens | $2.20 |
| Weights | Open |
| Parameters | 550B total / 55B active (MoE) |
| Reasoning levels | low, high, max |
| Modalities | text |
| API model id | nemotron-3-ultra-550b-a55b |
| Released | June 4, 2026 |
Pricing tiers: Open weights (weights + data + recipes); $0.50/$2.20 per MTok is the third-party hosted rate (OpenRouter). 550B/55B MoE.
Verified evidence
Published benchmark results
Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.
| Benchmark | Score | Measured | Source |
|---|---|---|---|
| Artificial Analysis Intelligence Index | 54 | 2026-08-14 | Artificial Analysis · View source |
Best for
- Open-weight frontier
- Self-hosting on NVIDIA GPUs
- Reasoning
Watch out
Source receipts
Catalog figures for this model were checked against the following sources.
- NVIDIA — Nemotron (accessed 2026-08-29)
- NVIDIA Nemotron 3 Ultra (accessed 2026-08-29)
Common questions
NVIDIA Nemotron 3 Ultra
Answered from the verified figures on this page rather than general guidance.
How much does NVIDIA Nemotron 3 Ultra cost per million tokens?
NVIDIA Nemotron 3 Ultra is listed at $0.50 per million input tokens and $2.20 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate. Open weights (weights + data + recipes); $0.50/$2.20 per MTok is the third-party hosted rate (OpenRouter). 550B/55B MoE.
What is NVIDIA Nemotron 3 Ultra's context window?
NVIDIA Nemotron 3 Ultra accepts about 1M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.
What is NVIDIA Nemotron 3 Ultra best for?
NVIDIA Nemotron 3 Ultra is a frontier tier from NVIDIA. It suits open-weight frontier, self-hosting on nvidia gpus, reasoning. Top US open-weight until Thinking Machines Inkling; verify hosted pricing.
Can I self-host NVIDIA Nemotron 3 Ultra?
NVIDIA Nemotron 3 Ultra publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.