Quick Facts
- Nemotron 3 Super delivers 5x higher throughput and 2x accuracy compared to previous generation
- 120 billion parameters with 12 billion active using mixture-of-experts architecture
- Early adopters include ServiceNow, Oracle, Palantir, and CrowdStrike for enterprise workflows
Nvidia released Nemotron 3 Super, its most advanced AI model designed for enterprise agentic systems. The model delivers five times higher throughput than its predecessor while doubling accuracy.
The 120-billion parameter model uses a hybrid mixture-of-experts architecture with 12 billion active parameters. It features a 1-million-token context window that allows AI agents to maintain full workflow state without losing focus.
Nemotron 3 Super scored 60.47% on the SWE-Bench Verified benchmark compared to GPT-OSS-120B’s 41.90%. The model achieved 91.75% on RULER at 1M tokens versus GPT-OSS’s 22.30%.
“Open innovation is the foundation of AI progress,” said Jensen Huang, Nvidia’s founder and CEO. “With Nemotron, we’re transforming advanced AI into an open platform that gives developers the transparency and efficiency they need to build agentic systems at scale.”
Major enterprises are already integrating Nemotron models into their operations. ServiceNow, Accenture, Deloitte, Oracle Cloud Infrastructure, and Zoom are using the technology for manufacturing, cybersecurity, software development, and communications workflows.
Bill McDermott, ServiceNow’s chairman and CEO, stated: “ServiceNow’s intelligent workflow automation combined with NVIDIA Nemotron 3 will continue to define the standard with unmatched efficiency, speed and accuracy.”
The model will be available through multiple cloud providers including Google Cloud’s Vertex AI and Oracle Cloud Infrastructure. Amazon Web Services and Microsoft Azure support is coming soon. Nvidia also partnered with Dell Technologies and Hewlett Packard Enterprise for enterprise access.
Nvidia optimized Nemotron 3 Super for its Blackwell GPU platform, achieving 4x faster inference than 8-bit models on the previous Hopper architecture. The company pre-trained the model in NVFP4 4-bit floating point format without accuracy loss.
The launch comes ahead of Nvidia’s GTC conference starting March 16. Nvidia plans to invest $26 billion over five years building open source AI models as it evolves from chipmaker to AI lab.
The model uses architectural innovations including Latent MoE that activates four times more expert specialists at the same inference cost. Multi-token prediction enables the system to predict multiple future tokens simultaneously, reducing generation time for long sequences.
Read more: Nvidia’s Nemotron Super 3 model for agentic systems launches with five times higher throughput
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
