Techday
This page was machine-translated and may differ from the original. View original

NVIDIA Innovates AI Training and Inference with NVFP4… Securing Both Performance and Efficiency

Google 우선 소스Published2026.02.23 16:35

Operate large-scale models without compromising quality while significantly reducing memory usage

NVIDIA is accelerating the paradigm shift in AI computing by dramatically boosting AI training and inference performance through NVFP4, a next-generation low-precision computing format.

NVIDIA announced on the 23rd that NVFP4-based technology enables the optimization of large-scale AI workloads by simultaneously improving throughput and energy efficiency while maintaining high accuracy.

As the scale and complexity of AI models increase rapidly, it is becoming difficult to meet performance demands relying solely on Moore's Law. In response, NVIDIA has presented a new solution through a co-design strategy that encompasses both hardware and software.

NVFP4, introduced with the Blackwell architecture, is based on 4-bit floating-point precision and provides higher computational density and energy efficiency compared to FP8.

The performance of NVFP4 has been clearly demonstrated in the latest MLPerf training benchmarks. The NVIDIA GB300 NVL72 system, composed of 512 Blackwell Ultra GPUs, completed the Llama 3.1 405B pre-training in 64.6 minutes, recording up to 1.9 times faster performance compared to the previous FP8-based generation.

In the inference domain as well, it significantly improved token throughput while maintaining high accuracy in major large language models such as DeepSeak-R1 and the Llama series.

NVFP4 demonstrates strengths even in long contexts and large-scale deployment environments. It enables the operation of large-scale models without compromising quality while drastically reducing memory usage, significantly improving the cost efficiency of AI services.

These characteristics are particularly advantageous for demanding workloads such as agentic AI, scientific simulations, and large-scale reinforcement learning.

Global companies such as Black Forest Labs, Radical Numerics, Cognition, and Red Hat are also participating in the NVFP4 ecosystem and contributing to the spread of the technology.

NVIDIA provides extensive support for NVFP4-based training and inference through various software stacks, including TensorRT-LLM, Torch.ao, and Transformer Engine.

The upcoming Rubin platform further enhances NVFP4 performance, heralding another leap forward in the speed and efficiency of AI training and inference.

NVFP4 is establishing itself as a core technology for next-generation AI infrastructure, accelerating the era of high-performance, low-cost AI.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
배종인 기자
배종인 기자