Nemotron 3 Ultra
by NVIDIAtext->text4 endpoints
Neura Intelligence Index
60.1/ 100
Rank
#132—
Confidence
mediumBenchmarks
5(13% coverage)
Overview
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Capabilities
Text generation
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling
Stop sequences
Deterministic seed
Modalities
Input
Text
Output
Text
Technical Specifications
Context Window
262.1K tokens
Max Output
32.8K tokens
Tokenizer
Other
Uptime (24h)
10000.0%
Parameters
550000000000B
License
openmdwOpen Weights
Country
🇺🇸US
Release Date
June 4, 20263mo
Supported Parameters
frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Benchmark Performance
coding
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Verified | 70.7 | verified |
| SciCode | 44.6 | verified |
reasoning
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 87 | verified |
| Humanity's Last Exam | 37.4 | verified |
agent
| Benchmark | Score | Source |
|---|---|---|
| BrowseComp | 44.4 | verified |
Performance Over Time
Pricing
Per 1M tokens
Input$0.63
Output$3.13
Cache read$0.19
Added
June 4, 2026
Last synced 9/18/2026