Tektronix TIF 2026
This page was machine-translated and may differ from the original. View original

NVIDIA Unveils GB200 NVL72, Delivering a 10x Improvement in AI Model Performance with MoE Architecture

Google 우선 소스Published2025.12.05 11:20

▲The Blackwell NVL72, with its ultra-collaborative design, is a game changer for MoE models.

72 Blackwell GPUs connected via NVLink to operate as a single system

NVIDIA, a leader in AI computing technology, has maximized the performance of the mixture-of-experts (MoE) model architecture adopted by cutting-edge AI models by 10x.

On the 4th, NVIDIA unveiled the technical achievements of the NVIDIA Blackwell GB200 NVL72.

This announcement further demonstrates NVIDIA's technological leadership in AI computing, significantly expanding its potential for use in global AI data centers and across key industries.

MoE is a structure that mimics the efficiency of the human brain, distributing tasks to specialized 'Experts' and activating only the necessary Experts for each token.

This method allows for faster and more efficient token generation without increasing computational complexity. All of the top 10 open-source models on the leaderboard of independent evaluation firm Artificial Analysis (AA) have adopted the MoE architecture, including DeepSeek-R1, Kimi K2 Thinking, and Mistral Large 3.

The NVIDIA GB200 NVL72 makes these MoE models scalable in real-world production environments. Specifically, the Kimi K2 thinking model achieved a 10x performance boost on NVL72 compared to the existing HGX H200, and the same performance was demonstrated on DeepSeak-R1 and Mistral Large 3.

Frontier-class MoE models are of a scale and complexity that is difficult to handle with a single GPU. The NVIDIA GB200 NVL72 provides a rack-scale architecture that connects 72 Blackwell GPUs via NVLink, operating as a single system. This supports 1.4 exaflops of AI performance, 30TB of shared memory, and 130TB per second of NVLink bandwidth.

This design solves existing bottlenecks by minimizing memory burden by reducing the number of experts per GPU and eliminating delays through ultra-high-speed communication between experts based on NVLink.

Additionally, the NVIDIA Dynamo framework is combined with open-source inference tools such as TensorRT-LLM, SGLang, and vLLM to maximize the inference performance of MoE models.

The GB200 NVL72 is being deployed through major cloud service providers including AWS, Google Cloud, Microsoft Azure, Oracle Cloud, CoreWeave, and Together AI.

“NVL72 is an integrated platform that delivers performance, scalability, and stability, delivering innovations only possible in a dedicated AI cloud,” said Peter Salanki, CTO of CoreWeave.

DeepL is also leveraging GB200 NVL72 to improve MoE model training and inference efficiency, and Fireworks AI deployed its Kimi K2 model on NVL72 to achieve the highest performance on the AA leaderboard.

“The GB200 NVL72 delivers a 10x performance improvement over Hopper in the DeepSeek-R1 model,” NVIDIA CEO Jensen Huang said at GTC Washington, D.C.

This results in a tenfold increase in token processing capacity, fundamentally changing the power and cost constraints of data centers.

Mistral Large 3 also achieved a 10x performance improvement over its predecessor in NVL72, demonstrating improved user experience, lower cost per token, and greater energy efficiency.

NVIDIA GB200 NVL72 provides optimized performance not only for MoE models but also for multimodal AI and agentic systems.

It efficiently supports diverse modalities, including language, visual, and audio, as well as agent-based workflows, and leverages a shared pool of experts to ensure scalability even in large-scale production environments.

NVIDIA is simultaneously realizing AI performance, efficiency, and scalability through the GB200 NVL72, and plans to further expand the possibilities of frontier models through the future Vera Rubin architecture roadmap.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
배종인 기자
배종인 기자