All posts

Models

NVIDIA releases Nemotron 3 Ultra, a 550B open-weights model built for long-running agents

NVIDIA published weights, data and recipes for a 550-billion-parameter hybrid Mamba-attention model. It computes like a 55B model but needs memory for all 550B.

HackHoster Team · · 11 min read

NVIDIA's headquarters in Santa Clara, California, with a large white lattice canopy over the buildings and an NVIDIA sign by the road
Photo: Coolcaesar / Wikimedia Commons, CC BY-SA 4.0

At a glance

  • NVIDIA put Nemotron 3 Ultra on Hugging Face on June 4, 2026, three days after Jensen Huang announced it at GTC Taipei during Computex.
  • The model has 550 billion parameters, but its router sends each token to 22 of 512 experts per layer, so about 55 billion do the work.
  • Artificial Analysis scored it 48 on its Intelligence Index, the top US open-weights result, but below Moonshot's Kimi K2.6 at 54.
  • The 4-bit NVFP4 checkpoint runs on a single node of four B200 GPUs; the BF16 version needs eight Blackwell GPUs or sixteen H100s.
  • NVIDIA also released 10 million new fine-tuning samples, 1 million new RL tasks and 173 billion tokens of fresh GitHub code.

NVIDIA has put the weights of Nemotron 3 Ultra on Hugging Face, along with most of the data and the training recipes behind it. Jensen Huang announced the model in his GTC Taipei keynote at Computex on June 1, and the checkpoints went live on June 4. The model is also on ModelScope, OpenRouter and build.nvidia.com, and NVIDIA ships it as a NIM microservice for its own deployment stack.

The headline numbers: 550 billion total parameters, about 55 billion active for each token, a context window of up to 1 million tokens, and roughly 20 trillion tokens of pretraining text. The license is OpenMDW-1.1, which the model card describes as allowing both commercial and non-commercial use. NVIDIA published three checkpoints: a base model, a post-trained BF16 model, and a 4-bit NVFP4 version meant for serving.

NVIDIA's pitch is that this is a model for agents, not for chat. Its announcement claims up to 5x faster inference and up to 30% lower cost on complex agentic tasks compared with leading open models, and lists early adopters including Perplexity, Palantir, ServiceNow, CrowdStrike, Glean, Harvey and CodeRabbit.

Jensen Huang on a large stage in front of a screen showing a rising green curve labelled perception AI, generative AI, agentic AI and physical AI
Jensen Huang at NVIDIA's CES 2025 keynote, where agentic AI was already on the slide. This is not the Computex 2026 keynote where he announced Nemotron 3 Ultra. Photo: Pronoia / Wikimedia Commons, CC0

Where Nemotron 3 Ultra comes from

Nemotron is NVIDIA's family of open models, and the third generation arrived in three sizes over six months. The smallest, Nemotron 3 Nano, came out on December 15, 2025. Super followed on March 11, 2026. Ultra completes the set.

All three share one design idea: replace most of a transformer's attention layers with Mamba-2 state-space layers, and replace dense feed-forward layers with mixture-of-experts layers. The Nano model card spells out what that looks like at small scale: 52 layers, of which 23 are Mamba-2, 23 are MoE and only 6 use attention. Each Nano MoE layer has 128 routed experts plus one shared expert, and 6 are active per token.

NVIDIA's March blog post on Super explains why it built the family this way. Agent systems produce far more tokens than ordinary chat, because every tool call, file read and intermediate plan goes back into the context. NVIDIA calls this context explosion and says multi-agent systems can generate up to 15 times the tokens of a standard conversation. A model that has to reread a growing transcript at every step gets slower and more expensive as the task goes on, and the company argues it also drifts from its original goal.

ModelReleasedTotal parametersActive per tokenPretraining tokensContext window
Nemotron 3 NanoDec 15, 202530B3.5B25T1M tokens
Nemotron 3 SuperMar 11, 2026120B12B25T1M tokens
Nemotron 3 UltraJun 4, 2026550B55Babout 20T1M tokens

Ultra is trained on fewer tokens than its smaller siblings. As described below, that was not the original plan.

The Taipei Nangang Exhibition Center, a curved glass and metal building, with a large green Computex 2016 banner and visitors walking in front
The Taipei Nangang Exhibition Center during Computex 2016. NVIDIA announced Nemotron 3 Ultra at its GTC Taipei event during Computex 2026. Photo: NVIDIA Taiwan / Wikimedia Commons, CC BY 2.0

How the architecture works

Nemotron 3 Ultra stacks four ideas on top of each other. Each one attacks a different cost.

Mamba-2 layers carry a running summary

A standard transformer runs attention in every layer, and attention compares each new token with every token before it. That work grows with the length of the context, and so does the key-value cache that has to stay in GPU memory. Mamba-2 layers are state-space layers. They read the sequence in order and carry a fixed-size running state forward, so, as NVIDIA's Super post puts it, their cost grows linearly with sequence length rather than quadratically. That is what makes a 1-million-token window practical.

A diagram of a recurrent network: a hidden state box h with a loop, unfolded into a chain of hidden states over time steps, each taking an input x and producing an output o
A recurrent network unfolded over time. Mamba-2 layers are not classic RNNs, but they share the key property shown here: a fixed-size state passed from one token to the next. Diagram: fdeloche / Wikimedia Commons, CC BY-SA 4.0

The trade-off is that a running summary is worse at pulling one exact fact out of a long document. NVIDIA keeps a small number of attention layers for that kind of precise recall and lets Mamba-2 handle the rest.

A few attention layers, with a tiny cache

According to MarkTechPost's summary of the architecture, Ultra has 108 layers and a model width of 8,192. Its attention layers use 64 query heads but only 2 key-value heads. Sharing keys and values across many query heads shrinks the cache that grows with context length, which matters a great deal when the context is a million tokens.

A block diagram of multi-headed attention: query, key and value inputs copied into several attention heads, concatenated, and passed through an output projection
Multi-headed attention, the building block Nemotron 3 Ultra keeps in only a few of its layers. Its attention layers share 2 key-value heads across 64 query heads. Diagram: Cosmia Nebula / Wikimedia Commons, CC BY-SA 4.0

LatentMoE routes tokens through a smaller space

Each MoE layer holds 512 small feed-forward networks, called experts, and a router sends every token to 22 of them. NVIDIA's version, LatentMoE, first projects each token into a smaller latent dimension and does the routing and expert computation there. NVIDIA's Super post says this lets the model consult four times as many experts for the same compute. MarkTechPost describes the trade as giving up some hidden-dimension width in exchange for more routed experts at a fixed inference cost.

Total versus active parameters. In a mixture-of-experts model, the total count includes every weight in every expert; the active count includes only the weights that run for a given token. Active parameters set the compute per token. Total parameters set how much memory you need, because any expert can be picked at any time.

A diagram of a mixture-of-experts layer: an input vector x feeds a gating network and eight expert boxes; the gate weights each expert's output before they are combined
A generic mixture-of-experts layer with a gating network and eight experts. Nemotron 3 Ultra has 512 experts per MoE layer and activates 22 for each token. Diagram: Sinafe / Wikimedia Commons, CC BY-SA 4.0

Multi-token prediction and 4-bit training

Ultra also predicts several future tokens per forward pass. MarkTechPost reports that its multi-token prediction heads share parameters during training, and the model card's serving examples use them for speculative decoding, drafting up to five tokens ahead that the main model then checks. NVIDIA's Super post credited this kind of drafting with up to 3x wall-clock speedups on structured output such as code and tool calls.

Finally, the model was pretrained in NVFP4, NVIDIA's 4-bit floating-point format. According to the model card, most linear layers use NVFP4 for weights, activations and gradients, while sensitive pieces such as latent projections, attention projections, embeddings and the prediction heads stay in BF16 or MXFP8. MarkTechPost describes the format as E2M1, meaning one sign bit, two exponent bits and one mantissa bit, with scaling factors shared across two-dimensional blocks of values to keep the tiny numbers in range.

The bit layout of a 32-bit floating-point number: one sign bit, eight exponent bits and twenty-three fraction bits
The 32-bit IEEE 754 layout, for comparison. NVFP4 squeezes each value into 4 bits, one sign, two exponent and one mantissa bit, plus shared scale factors per block. Diagram: Codekaizen / Wikimedia Commons, CC BY 3.0

How NVIDIA trained it

Pretraining ran for about 20 trillion tokens in two phases, according to MarkTechPost: 15 trillion tokens weighted toward diversity, then 5 trillion weighted toward quality, followed by an extension of the context window to 1 million tokens. The model card gives a pretraining data cutoff of September 2025 and a post-training cutoff of May 2026.

The run did not go smoothly, and NVIDIA's account of it is one of the more useful parts of the release for anyone training large models. MarkTechPost reports two loss divergences. Near 8 trillion tokens, a change had moved the gradient reduction for the output layer from FP32 to BF16, and the gradient from the prediction heads was getting lost in BF16's 7 mantissa bits. Reverting to FP32 fixed it. Near 16 trillion tokens, a second divergence had no confirmed root cause. NVIDIA annealed the learning rate early and cut the planned token horizon to 20 trillion, which is why Ultra saw fewer tokens than Nano or Super.

Post-training ran in stages: supervised fine-tuning, reinforcement learning with verifiable rewards, and then what NVIDIA calls multi-domain on-policy distillation, in which more than ten domain-specific teacher models guide training on the student's own attempts rather than on offline examples. MarkTechPost says NVIDIA ran two rounds of distillation, re-initializing the teachers from the improved student in between, and added 15 new RL environments.

The data release is substantial. This round adds 10 million new supervised fine-tuning samples, for 50 million in total, 1 million new RL tasks, for 2 million in total, and 173 billion tokens of refreshed GitHub code with a cutoff of September 30, 2025. The model card points to two Hugging Face collections, Nemotron-Pre-Training-Datasets and Nemotron-Post-Training-v3, and notes that some sets require approval because of their source licenses.

The numbers

NVIDIA's model card reports results for both the BF16 and the 4-bit checkpoints. A selection:

BenchmarkWhat it measuresBF16NVFP4
Terminal Bench 2.1Agent tasks in a command-line sandbox56.453.9
PinchBenchAgentic task completion90.089.8
τ-bench v3, averageTool-using customer-service agents70.970.3
τ-bench v3, bankingThe hardest τ-bench domain listed22.619.2
BrowseCompWeb research with search44.441.4
GPQA, no toolsGraduate-level science questions87.087.9
HLE, no toolsHumanity's Last Exam26.726.1
IOI 2025Competitive programming score570.0564.7
RULER at 1M tokensLong-context retrieval94.794.0

Two things stand out. The 4-bit checkpoint loses little, and on a few tests it scores slightly higher, which supports NVIDIA's case that NVFP4 is usable for serving rather than a lossy shortcut. And the τ-bench banking score is far below the other domains, a reminder that a strong average can hide a weak spot that matters for a specific product.

On speed, MarkTechPost reports NVIDIA's own throughput comparisons with 8K input and 64K output tokens on GB200 hardware: 5.9x the throughput of GLM-5.1, 4.8x Kimi K2.6 and 1.6x Qwen 3.5. The same report notes that the comparison used TensorRT-LLM for Nemotron and vLLM for GLM, so some of the gap is the serving stack, and that Ultra trails Qwen 3.5 on prefill-heavy work with 50K input and 2K output tokens. NVIDIA's 30% cost claim comes from Ultra using fewer tokens per turn on SWE-Bench and Terminal Bench.

Caveat on the scores. Apart from Artificial Analysis's index, every benchmark number in this post comes from NVIDIA, and some evaluations used internal scaffolding the model card says NVIDIA plans to open-source later. Treat them as a starting point for your own tests.

What outside testers found

Artificial Analysis, which runs its own benchmark suite, evaluated the model ahead of launch. It scored Nemotron 3 Ultra 48 on its Intelligence Index, the highest of any US open-weights model it has measured. Google's Gemma 4 31B scored 39, NVIDIA's own Nemotron 3 Super 36 and OpenAI's gpt-oss-120b 33. Ultra is not the top open model overall: Moonshot's Kimi K2.6, from China, scores 54.

The same group measured more than 300 output tokens per second on a pre-release DeepInfra endpoint, against 50 to 100 for the Chinese models it compared. That supports the speed claim on at least one provider.

CodeRabbit, the code-review company and one of NVIDIA's launch partners, published a more mixed picture. On its set of 105 review problems, Ultra passed 58 against 60 for CodeRabbit's unnamed baseline models, with a mean latency of 7:06 against 8:31 for the baseline. The problem was reliability: CodeRabbit counted an average of 36.5 retries for the Ultra runs, compared with 0.3 for the baseline, because the model sometimes stopped before producing the required structured output. Retrying the same prompt usually worked, but CodeRabbit wrote that "the first-attempt completion behavior is not stable enough to ignore."

That finding matters more for agent builders than any single benchmark. A model that is fast but needs a retry loop around every structured step shifts cost and complexity into your harness.

What self-hosting takes

MoE saves compute, not memory. All 550 billion parameters have to be loaded, because any token can be routed to any expert. At 16 bits per weight that is about 1.1 TB before you store a single token of context.

An NVIDIA HGX B200 board with eight GPU modules under large black heatsinks, sitting on a wooden table
An HGX B200 board carries eight Blackwell GPUs. NVIDIA recommends four B200s for the 4-bit checkpoint and eight for full BF16 weights. Photo: Pokiiri / Wikimedia Commons, CC BY-SA 4.0

NVIDIA's two model cards set the hardware floor:

CheckpointMinimum hardware NVIDIA listsNotes
BF168x GB200, B200, GB300 or B300; 8x H200; or 16x H100Single node of 8 B200s recommended
NVFP44x B200 on one node, or at least 4 GPUs across GB200 or GB300 systemsHopper runs it with 4-bit weights and 16-bit activations

The 4-bit checkpoint mixes precisions. MarkTechPost reports that routed experts are stored in NVFP4, shared and Mamba layers in FP8, and attention layers in BF16, for an average of 5.03 bits per weight. At that rate the weights take roughly 350 GB, about a third of the BF16 size. Blackwell GPUs do the 4-bit math natively; Hopper GPUs, which lack FP4 units, fall back to 4-bit weights with 16-bit activations.

The cards document serving with vLLM, SGLang and TensorRT-LLM. The default examples cap context at 256K tokens and require an environment flag to go to the full million, a sign of how much cache memory the long window costs even with Mamba layers.

What builders can do with it

For a hackathon team, hosted APIs are the realistic way in. OpenRouter and build.nvidia.com let you drop the model into an agent loop without renting a GPU node, and MarkTechPost lists Nebius, Together AI and Perplexity among the inference partners.

A few things worth trying:

  • Use the reasoning controls. The chat template has an enable_thinking flag and a medium-effort mode. MarkTechPost reports that medium effort uses about 2.5 times fewer reasoning tokens for about a 7% accuracy cost, which can be the better trade for an agent that makes many small calls.
  • Build the retry loop first. Given CodeRabbit's experience, validate every structured output and retry on failure. The model card also recommends a force_nonempty_content setting for coding agents.
  • Test your hardest domain, not the average. The banking result above shows how uneven agent scores can be. Run your own task set before switching.
  • Mine the data even if you skip the model. The released SFT samples and RL tasks show how a large agent model was post-trained, and you can reuse them to fine-tune a smaller open model on hardware you can afford, including Nemotron 3 Nano.

Practical tip. If your agent spends most of its time reading long inputs rather than writing long outputs, benchmark against Qwen 3.5 too. NVIDIA's own numbers show Ultra's speed lead shrinks or reverses on prefill-heavy workloads.

Teams with cloud credits for an eight-GPU Blackwell or H200 machine can run vLLM or the NIM container and keep their data in-house. The model card lists ten supported languages: English, French, Spanish, Italian, German, Japanese, Korean, Hindi, Brazilian Portuguese and Chinese.

Open questions and what to watch

The biggest open question is whether speed beats raw quality for agent work. On Artificial Analysis's index, the best Chinese open models still score higher. NVIDIA's bet is that an agent which finishes a task in fewer, faster, cheaper steps is worth more than a few index points, and that a US-origin model with published data will matter to some enterprise buyers. Both are plausible, and neither is proven by the launch numbers.

Long context is the second. RULER at 1 million tokens tests whether a model can find planted facts, not whether an agent stays on task across a long session. Filling a million-token window also still costs memory and money, even with most attention layers replaced.

As of June 5, the things to watch are concrete: independent agent evaluations beyond Artificial Analysis's index, whether the retry behavior CodeRabbit saw improves in later checkpoints or serving templates, and how quickly the vLLM and SGLang integrations mature for the 4-bit checkpoint on Hopper hardware, which is what most teams outside the largest labs can actually rent.

Sources