Skip to main content
AI Choice EngineAI Choice Engine

NVIDIA

NVIDIA Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is a NVIDIA model with 1M context, $0.50 in / $2.20 out per million tokens, last verified September 2, 2026. Open weights: yes. Public-board scores are linked to the publisher; vendor-only figures stay in the watch-out, not on the board.

Nemotron 3 Ultra is NVIDIA's open-weight (550B/55B) frontier MoE with a 1M-token context.

Catalog record checked September 2, 2026Individual provider fields may change

AI model specification details
SpecificationNVIDIA Nemotron 3 Ultra
ProviderNVIDIA
TierFrontier
Context window1M
Max outputNot verifiedUnverified
Input / 1M tokens$0.50
Output / 1M tokens$2.20
WeightsOpen
Parameters550B total / 55B active (MoE)
Reasoning levelslow, high, max
Modalitiestext
API model idnemotron-3-ultra-550b-a55b
ReleasedJune 4, 2026

Pricing tiers: Open weights (weights + data + recipes); $0.50/$2.20 per MTok is the third-party hosted rate (OpenRouter). 550B/55B MoE.

Verified evidence

Published benchmark results

Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.

Published benchmark results for NVIDIA Nemotron 3 Ultra
BenchmarkScoreMeasuredSource
Artificial Analysis Intelligence Index542026-08-14Artificial Analysis · View source

Best for

  • Open-weight frontier
  • Self-hosting on NVIDIA GPUs
  • Reasoning

Watch out

Top US open-weight until Thinking Machines Inkling; verify hosted pricing.

Source receipts

Catalog figures for this model were checked against the following sources.

Common questions

NVIDIA Nemotron 3 Ultra

Answered from the verified figures on this page rather than general guidance.

How much does NVIDIA Nemotron 3 Ultra cost per million tokens?

NVIDIA Nemotron 3 Ultra is listed at $0.50 per million input tokens and $2.20 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate. Open weights (weights + data + recipes); $0.50/$2.20 per MTok is the third-party hosted rate (OpenRouter). 550B/55B MoE.

What is NVIDIA Nemotron 3 Ultra's context window?

NVIDIA Nemotron 3 Ultra accepts about 1M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.

What is NVIDIA Nemotron 3 Ultra best for?

NVIDIA Nemotron 3 Ultra is a frontier tier from NVIDIA. It suits open-weight frontier, self-hosting on nvidia gpus, reasoning. Top US open-weight until Thinking Machines Inkling; verify hosted pricing.

Can I self-host NVIDIA Nemotron 3 Ultra?

NVIDIA Nemotron 3 Ultra publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.