xpu liveBETA
Models›Nemotron 3 Ultra

Nemotron 3 Ultra

Model reference ↗

Provider prices 2

Lowest input price per billing unit. Conditions in Details.

ProviderPriceProvider model IDStatusContext / max outputSource
OpenRouterFree variant
Router catalog
Input $0USD / 1M tokensOutput $0USD / 1M tokens
nvidia/nemotron-3-ultra-550b-a55b:freeCatalog-listed
Conditions

Check regional and account eligibility in provider documentation.

Context tokens 1000000Output tokens 65536Official source ↗2026-09-28
DeepInfra
Standard list
Input $0.5USD / 1M tokensOutput $2.2USD / 1M tokens
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55BCatalog-listed
Conditions

Catalog-listed; account access and regional availability may apply.

Context tokens 262144Official source ↗2026-09-28
OpenRouter
Router catalog
Input $0.5USD / 1M tokensOutput $2.2USD / 1M tokens
nvidia/nemotron-3-ultra-550b-a55bCatalog-listed
Conditions

Catalog-listed; account access and regional availability may apply.

Context tokens 262144Output tokens 182520Official source ↗2026-09-28

Nemotron 3 Ultra accepts text and produces text. Tool calling and structured outputs are available on supported routes.

CONTEXT TOKENS1,048,576Developer model card
MAX OUTPUT TOKENS182,520OpenRouter catalog
PROVIDERS2Linked in this catalog
SOURCE CHECKED2026-09-29Sources reviewed; unknown fields labeled

Specifications

Catalog model ID
nvidia/nemotron-3-ultra-550b-a55b
Developer model ID / revision
Not verified in cited sources
Input modalities
textOpenRouter catalog ↗
Output modalities
textOpenRouter catalog ↗
Context / input token limit
1,048,576Developer model card ↗
Maximum output tokens
182,520OpenRouter catalog ↗
Tool calling
SupportedOpenRouter catalog ↗
Structured output
SupportedOpenRouter catalog ↗
API parameters (source-specific)
frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_pOpenRouter catalog ↗
API parameter scope
OpenRouter API; support varies by route
Weight access
Open weightsDeveloper model card ↗
Weight license / access terms
OpenMDW 1.1Developer model card ↗
Total parameters (billions)
550Developer model card ↗
Active parameters (billions)
55Developer model card ↗
Download size (GB)
Depends on checkpoint and precision; no verified file size
Release date
2026-06-04Developer model card ↗
Added to OpenRouter
2026-06-04OpenRouter catalog ↗
Notes
API parameter names are reported by OpenRouter and do not describe the developer’s native API. Token limits labeled OpenRouter are route catalog limits, not universal model limits. Media limits on this page describe the linked provider endpoint; other deployments may differ. The model card supports up to 1M context; the OpenRouter route currently exposes 262,144 tokens. Check deployment settings before using the full model window.

Catalog entries link to their specification and availability sources. This is a curated selection, not every model each provider offers. Best prices are the lowest collected USD tariffs, not a total-cost estimate. Token offers are selected by input price, with output from the same tariff; workload mix may change the cheapest provider. Listed variants may require batch processing, off-peak hours, or specific media settings. Price coverage varies by provider; collected tariffs include their source and observation time. Benchmark scores are a dated snapshot, not a live feed or a universal measure of quality. They apply to the shown evaluation configuration; listed prices may use different settings. Models without verified scores are not assumed to be weaker.