Nvidia Unveils Nemotron 3 Super, a 120B Open Model Built to Cut AI Agent Inference Costs

Nvidia Unveils Nemotron 3 Super, a 120B Open Model Built to Cut AI Agent Inference Costs

N
News Editor 01
2026-07-23 03:15:14
Nvidia has launched Nemotron 3 Super, a 120B open hybrid model that activates only 12.7B parameters per forward pass, targeting lower-cost and higher-throughput inference for AI agent workloads.
NvidiaAI agentsopen modellarge language modelNemotron 3 Super

Nvidia has introduced Nemotron 3 Super, a 120 billion-parameter open hybrid model aimed at AI agent workloads. The company said the model uses a Mixture-of-Experts architecture and activates only 12.7 billion parameters per forward pass, a design intended to reduce the compute burden of running agents at scale. The release targets a practical problem for developers: multi-step reasoning and multi-agent systems can drive token consumption sharply higher, pushing inference costs up with it.

According to Nvidia, token usage in multi-agent pipelines can expand by as much as 15x. Nemotron 3 Super is positioned as a response to that issue. It is the second model in the Nemotron 3 family, following Nemotron 3 Nano, which was released in December 2025. Nvidia said the new model was announced around March 10, 2026.

A hybrid 88-layer backbone built for 1 million-token context

Nemotron 3 Super is built on a hybrid Mamba-Transformer backbone spanning 88 layers. Mamba-2 blocks are used for long-sequence processing with linear-time efficiency, while Transformer attention layers are kept for precise recall. Nvidia said that combination gives the model native support for context windows up to 1 million tokens without the memory overhead usually associated with pure-attention systems.

The model also includes a LatentMoE routing system. Token embeddings are compressed into a low-rank space before being sent to 512 experts per layer, with 22 experts activated at a time. Nvidia said this setup allows roughly 4x more experts at the same inference cost than standard MoE methods, while supporting narrower specialization, including separating Python logic from SQL handling at the expert level.

Trained on 25 trillion tokens, with up to 3x faster generation on structured tasks

Nemotron 3 Super uses Multi-Token Prediction layers with two shared-weight heads. Nvidia said these layers speed up chain-of-thought generation and enable native speculative decoding. On structured tasks, the company reported generation speeds of up to 3x faster.

Pre-training was carried out in two stages and totaled 25 trillion tokens. The first stage used 20 trillion tokens of broad data, and the second used 5 trillion high-quality tokens tuned for benchmark performance. A later extension phase on 51 billion tokens pushed native context length to 1 million tokens. Post-training included roughly 7 million supervised fine-tuning samples and reinforcement learning across 21 environments with more than 1.2 million rollouts.

Up to 7.5x the throughput of Qwen3.5-122B-A10B

On benchmarks, Nemotron 3 Super scored 83.73 on MMLU-Pro, 90.21 on AIME25, and 60.47 on SWE-Bench with OpenHands. It reached 85.6% on PinchBench, which Nvidia described as the highest reported score among open models in its class. On the long-context test RULER 1M, it posted 91.64.

For throughput, Nvidia said that at 8k input and 64k output, Nemotron 3 Super delivers 2.2x the throughput of GPT-OSS-120B and up to 7.5x that of Qwen3.5-122B-A10B. Against the prior Nemotron Super generation, the company reported more than 5x throughput and as much as 2x the accuracy.

Open checkpoints, training data, and broad deployment options

Nvidia said the model was trained end-to-end in its NVFP4 four-bit floating-point format and optimized for Blackwell GPUs. On B200 hardware, inference can run up to 4x faster than FP8 on H100, with no reported accuracy loss. Quantized FP8 and NVFP4 checkpoints retain at least 99.8% of full-precision accuracy, according to the company.

Nemotron 3 Super is available under the Nvidia Nemotron Open Model License. Checkpoints in BF16, FP8, and NVFP4 formats, along with pre-training data, post-training samples, and reinforcement learning environments, have been published on Hugging Face. Inference access is available through Nvidia NIM, build.nvidia.com, Perplexity, Openrouter, Together AI, Google Cloud, AWS, Azure, and Coreweave, while on-premises deployment is supported through Dell Enterprise Hub and HPE. Nvidia also said developers can use the NeMo platform for training recipes, fine-tuning guides, and inference cookbooks with vLLM, SGLang, and TensorRT-LLM.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.