ai/nemotron-3-super

Verified Publisher

By Docker

Updated 6 days ago

Artifact
0

6.7K

ai/nemotron-3-super repository overview

Read our How to Run Nemotron 3 Super Guide!

See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks.

  • Note <think> and </think> are separate tokens, so use --special if needed.
  • You can also fine-tune the model with Unsloth.

NVIDIA-Nemotron-3-Super-120B-A12B-BF16

Model Summary

Total Parameters120B (12B active)
ArchitectureLatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP)
Context LengthUp to 1M tokens
Minimum GPU Requirement8× H100-80GB
Supported LanguagesEnglish, French, German, Italian, Japanese, Spanish, Chinese
Best ForAgentic workflows, long-context reasoning, high-volume workloads (e.g. IT ticket automation), tool use, RAG
Reasoning ModeConfigurable on/off via chat template (enable_thinking=True/False)
LicenseNVIDIA Nemotron Open Model License
Release DateMarch 11, 2026

Quick Start

Use temperature=1.0 and top_p=0.95 across all tasks and serving backends — reasoning, tool calling, and general chat alike.

For more details on how to deploy and use the model - see the Quick Start Guide below!

For running Nemotron 3 Super on a single B200 or DGX Spark - please see: NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

Model Overview

Model Developer: NVIDIA Corporation

Model Dates: December 2025 - March 2026

Data Freshness:

  • The post-training data has a cutoff date of February 2026.
  • The pre-training data has a cutoff date of June 2025.
What is Nemotron?

NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.

Description

Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template.

The model employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. Distinct from the Nano model, the Super model incorporates Multi-Token Prediction (MTP) layers for faster text generation and improved quality, and it is trained using NVFP4 quantization to maximize compute efficiency. The model has 12B active parameters and 120B parameters in total.

The supported languages include: English, French, German, Italian, Japanese, Spanish, and Chinese

This model is ready for commercial use.

License/Terms of Use

Governing Download Terms: Use of this model is governed by the NVIDIA Nemotron Open Model License.

Governing Download Terms with NIM: The NIM container is governed by the NVIDIA Software License Agreement and Product-Specific Terms for AI Products. Use of this model is governed by the NVIDIA Nemotron Open Model License.

Benchmarks
BenchmarkNemotron 3 SuperQwen3.5-122B-A10BGPT-OSS-120B
General Knowledge
MMLU-Pro83.7386.7081.00
Reasoning
AIME25 (no tools)90.2190.3692.50
HMMT Feb25 (no tools)93.6791.4090.00
HMMT Feb25 (with tools)94.7389.55
GPQA (no tools)79.2386.6080.10
GPQA (with tools)82.7080.09
LiveCodeBench (v5 2024-07↔2024-12)81.1978.9388.00
SciCode (subtask)42.0542.0039.00
HLE (no tools)18.2625.3014.90
HLE (with tools)22.8219.0
Agentic
Terminal Bench (hard subset)25.7826.8024.00
Terminal Bench Core 2.031.0037.5018.70
SWE-Bench (OpenHands)60.4766.4041.9
SWE-Bench (OpenCode)59.2067.40
SWE-Bench (Codex)53.7361.20
SWE-Bench Multilingual (OpenHands)45.7830.80
TauBench V2
    Airline56.2566.049.2
    Retail62.8362.667.80
    Telecom64.3695.0066.00
    Average61.1574.5361.0
BrowseComp with Search31.2833.89
BIRD Bench41.8038.25
Chat & Instruction Following
IFBench (prompt)72.5673.7768.32
Scale AI Multi-Challenge55.2361.5058.29
Arena-Hard-V273.8875.1590.26
Long Context
AA-LCR58.3166.9051.00
RULER @ 256k96.3096.7452.30
RULER @ 512k95.6795.9546.70
RULER @ 1M91.7591.3322.30
Multilingual
MMLU-ProX (avg over langs)79.3685.0676.59
WMT24++ (en→xx)86.6787.8488.89

All evaluation results were collected via Nemo Evaluator SDK and for most benchmarks, the Nemo Skills Harness. For reproducibility purposes, more details on the evaluation settings can be found in the Nemo Evaluator SDK configs folder and the reproducibility tutorial for Nemotron 3 Super. The open source container on Nemo Skills packaged via NVIDIA's Nemo Evaluator SDK used for evaluations can be found here. In addition to Nemo Skills, the evaluations also used dedicated open-source packaged containers for Tau-2 Bench (default prompt), Terminal Bench Hard (48 tasks), ScaleAI Multi Challenge Multi-turn Instruction Following, and Ruler.

The following benchmarks are not onboarded yet in our open source tools and for these we used either their official open source implementation or otherwise an internal scaffolding that we plan to open source in the future: SWE Bench Verified (OpenHands), SWE Bench Multilingual (OpenHands), BrowseComp with Search (internal implementation with Serp API), Terminal Bench Core 2.0 (Harbor).

Deployment Geography: Global
Use Case

NVIDIA-Nemotron-3-Super-120B-A12B-BF16 is a general purpose reasoning and chat model intended to be used in English, Code, and supported multilingual contexts. This model is optimized for collaborative agents and high-volume workloads. It is intended to be used by developers designing AI Agent systems, chatbots, RAG systems, and other AI-powered applications. This model is also suitable for complex instruction-following tasks and long-context reasoning.

Release Date

Hugging Face - 03/11/2026 via Hugging Face

Reference(s)

Model Architecture

  • Architecture Type: Mamba2-Transformer Hybrid Latent Mixture of Experts (LatentMoE) with Multi-Token Prediction (MTP)
  • Network Architecture: Nemotron Hybrid LatentMoE
  • Number of model parameters: 120B Total / 12B Active

Model Design

The model utilizes the LatentMoE architecture, where tokens are projected into a smaller latent dimension for expert routing and computation, improving accuracy per byte. The Super model is pre-trained using NVFP4 quantization — the first model in the Nemotron 3 family trained at this precision. The majority of linear layers use NVFP4 for weights, activations, and gradients, while select layers (including latent projections, MTP layers, QKV/attention projections, and embeddings) are maintained in BF16 or MXFP8 for training stability. The model includes Multi-Token Prediction (MTP) layers using a shared-weight design across prediction heads. This improves training signal quality, enables faster inference via native speculative decoding, and supports more stable autoregressive drafting at longer draft lengths compared to independently trained offset heads.

Training Methodology

Stage 1: Pre-Training

Stage 2: Supervised Fine-Tuning

  • The model was further fine-tuned on synthetic code, math, science, tool calling, instruction following, structured outputs, and general knowledge data. This stage incorporated data designed to support long-range retrieval and multi-document aggregation. All datasets are disclosed in the Training and Evaluation Datasets section of this document. Major portions of the fine-tuning corpus are released in the Nemotron-Post-Training-v3 collection. Data Designer is one of the libraries used to prepare these corpora.

Stage 3: Reinforcement Learning

  • The model underwent multi-environment reinforcement learning using asynchronous GRPO (Group Relative Policy Optimization) across math, code, science, instruction following, multi-step tool use, multi-turn conversations, and structured output environments. It utilized an asynchronous RL architecture that fully decouples training from inference across separate GPU devices, leveraging in-flight weight updates and MTP to accelerate rollout generation. Conversational quality was further refined through RLHF. All datasets are disclosed in the Training and Evaluation Datasets section of this document. The RL environments and datasets are released as part of NeMo Gym.
  • Software used for reinforcement learning: NeMo RL, NeMo Gym

NVIDIA-Nemotron-3-Super-120B-A12B-BF16 model is a result of the above work.

The end-to-end training recipe is available in the NVIDIA Nemotron Developer Repository. Evaluation results can be replicated using the NeMo Evaluator SDK. Data Designer is one of the libraries used to prepare the pre and post training datasets. More details on the datasets and synthetic data generation methods can be found in the technical report NVIDIA Nemotron 3 Super Technical Report.

Input

  • Input Type(s): Text
  • Input Format(s): String
  • Input Parameters: One-Dimensional (1D): Sequences
  • Other Properties Related to Input: Maximum context length up to 1M tokens. Supported languages include: English, French, German, Italian, Japanese, Spanish, and Chinese

Output

  • Output Type(s): Text
  • Output Format: String
  • Output Parameters: One-Dimensional (1D): Sequences
  • Other Properties Related to Output: Maximum context length up to 1M tokens

Our AI models are designed and optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.

Software Integration

  • Runtime Engine(s): NeMo 25.11.01
  • Supported Hardware Microarchitecture Compatibility: NVIDIA Ampere - A100; NVIDIA Blackwell; NVIDIA Hopper - H100-80GB
  • Operating System(s): Linux

The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.

Model Version(s)

  • v1.0 - GA

Quick Start Guide

For each inference backend - we'll be using the custom super_v3 reasoning parser - which you can obtain by following these instructions:

wget https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/main/super_v3_reasoning_parser.py

OR

curl -O https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/main/super_v3_reasoning_parser.py

For advanced deployment configurations - visit this resource

vLLM

For more detailed information, please see this cookbook.

pip install -U vllm --extra-index-url https://wheels.vllm.ai/097eb544e9a22810c9b7a59e586b61627b308362

export MODEL_CKPT=PATH/TO/MODEL/CHECKPOINT
vllm serve $MODEL_CKPT \
  --served-model-name nvidia/nemotron-3-super \
  --async-scheduling \
  --dtype auto \
  --kv-cache-dtype fp8 \
  --tensor-parallel-size 4 \
  --pipeline-parallel-size 1 \
  --data-parallel-size 2 \
  --max-model-len 262144 \
  --enable-expert-parallel \
  --attention-backend TRITON_ATTN \
  --swap-space 0 \
  --trust-remote-code \
  --gpu-memory-utilization 0.9 \
  --enable-chunked-prefill \
  --mamba-ssm-cache-dtype float16 \
  --reasoning-parser-plugin super_v3_reasoning_parser.py \
  --reasoning-parser super_v3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

Context length defaults to 256k above. To use up to 1M, set VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 and --max-model-len 1M

SGLang

Container:

docker pull lmsysorg/sglang:v0.5.9

Or pip:

pip install 'git+https://github.com/sgl-project/sglang.git#subdirectory=python'

For more detailed information, please see this cookbook.

python3 -m sglang.launch_server \
  --model PATH/TO/CHECKPOINT \
  --served-model-name nvidia/nemotron-3-super \
  --trust-remote-code \
  --tp 8 \
  --ep 4 \
  --tool-call-parser qwen3_coder \
  --reasoning-parser nano_v3

Context length defaults to 256k above. To use up to 1M, set SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 and --context-length 1048576

TRT-LLM

Container:

docker pull nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc5

For more detailed information, please see this cookbook.

cat > ./extra-llm-api-config.yml << EOF
kv_cache_config:
  enable_block_reuse: false
  mamba_ssm_cache_dtype: float32
moe_config:
  backend: TRTLLM
cuda_graph_config:
  enable_padding: true
  max_batch_size: 256
enable_attention_dp: true
EOF

trtllm-serve PATH/TO/BF16/CHECKPOINT \
  --host 0.0.0.0 \
  --port 8123 \
  --backend pytorch \
  --max_batch_size 256 \
  --tp_size 8 --ep_size 8 \
  --max_num_tokens 8576 \
  --trust_remote_code \
  --reasoning_parser nano_v3 \
  --tool_parser qwen3_coder \
  --extra_llm_api_options extra-llm-api-config.yml
API Client

The examples below use the OpenAI-compatible client and work with any of the serving backends above.

NOTE: For coding agents add the following to the API call - extra_body={“chat_template_kwargs”: {“force_nonempty_content”: True}

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
MODEL = "nvidia/nemotron-3-super"

Reasoning ON (default)

response = client.chat.completions.create(
    model=MODEL,
    messages=[{"role": "user", "content": "Write a haiku about GPUs"}],
    max_tokens=16000,
    temperature=1.0,
    top_p=0.95,
    extra_body={"chat_template_kwargs": {"enable_thinking": True}}
)
print(response.choices[0].message.content)

Reasoning OFF

response = client.chat.completions.create(
    model=MODEL,
    messages=[{"role": "user", "content": "What is the capital of Japan?"}],
    max_tokens=16000,
    temperature=1.0,
    top_p=0.95,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}}
)
print(response.choices[0].message.content)

Low-effort reasoning

Uses significantly fewer reasoning tokens than full thinking mode. Recommended as a starting point before tuning explicit token budgets.

response = client.chat.completions.create(
    model=MODEL,
    messages=[{"role": "user", "content": "What is the capital of Japan?"}],
    max_tokens=16000,
    temperature=1.0,
    top_p=0.95,
    extra_body={"chat_template_kwargs": {"enable_thinking": True, "low_effort": True}}
)
print(response.choices[0].message.content)
OpenCode

OpenCode is an AI coding agent that runs in your terminal. It connects to any OpenAI-compatible endpoint, making it compatible with all three serving backends above (vLLM, SGLang, and TRT-LLM).

Create or update your ~/.config/opencode/opencode.json:

{
    "$schema": "https://opencode.ai/config.json",
    "model": "local/nvidia-nemotron-3-super",
    "provider": {
        "local": {
            "npm": "@ai-sdk/openai-compatible",
            "name": "local_backend",
            "options": {
                "baseURL": "http://localhost:8000/v1",
                "apiKey": "EMPTY"
            },
            "models": {
                "nvidia-nemotron-3-super": {
                    "name": "nvidia/nemotron-3-super",
                    "limit": {
                        "context": 1000000,
                        "output": 32768
                    }
                }
            }
        }
    },
    "agent": {
        "build": {
            "temperature": 1.0,
            "top_p": 0.95,
            "max_tokens": 32000
        },
        "plan": {
            "temperature": 1.0,
            "top_p": 0.95,
            "max_tokens": 32000
        }
    }
}

Update baseURL to match whichever backend you are running. The default port above (8000) matches the vLLM example; SGLang and TRT-LLM use 30000 and 8123 respectively.

To learn more about other supported agent scaffolds - check out this resource

Advanced: Budget-Controlled Reasoning

Set a hard token ceiling on the reasoning trace using reasoning_budget. The model will attempt to close the trace at the next newline before the budget is hit; if none is found within 500 tokens it closes abruptly at reasoning_budget + 500.

from typing import Any, Dict, List
import openai
from transformers import AutoTokenizer


class ThinkingBudgetClient:
    def __init__(self, base_url: str, api_key: str, tokenizer_name_or_path: str):
        self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_name_or_path)
        self.client = ope

…(truncated — see the full README on HuggingFace)

Tag summary

Content type

Unrecognized

Digest

sha256:6976a8705

Size

76.9 GB

Last updated

6 days ago

docker pull ai/nemotron-3-super

This week's pulls

Pulls:

4,375

Last week