This page was machine-translated and may differ from the original. View original

NVIDIA Accelerates the Era of Personalized AI, Unveiling the Nemotron 3 Open Model

Google 우선 소스Published2025.12.16 11:47

Nemotron 3, based on a hybrid MoE architecture, simultaneously secures efficiency and accuracy.
Unsloth, LLM fine-tuning performance improved by 2.5x, optimized for NVIDIA hardware.

NVIDIA, a global leader in AI computing, accelerates the era of personalized AI by unveiling its new open model product line, Nemotron 3.

NVIDIA announced on the 16th that it has released Nemotron 3 and supports fine-tuning of large language models (LLMs) more quickly and efficiently through the open-source framework Unsloth.

This announcement is significant because it enables the building of customized AI assistants optimized for learning, work, and creative projects, based on the RTX AI PC and the ultra-compact supercomputer DGX Spark.

Nemotron 3 comes in three models: Nano, Super, and Ultra.

All are designed based on a hybrid Mixture-of-Experts (MoE) architecture to ensure both efficiency and accuracy.

The Nemotron 3 Nano 30B-A3B is our most compute-efficient model, optimized for software debugging, content summarization, and AI assistant workflows.

Reduce costs by up to 60% with inference tokens and 1 million token contexts It supports Windows and maintains high accuracy even during long-term, multi-step tasks.

Nemotron 3 Super is a high-precision inference model for multi-agent applications, and Ultra is a model for complex AI applications, and is scheduled to be released in the first half of 2026.

NVIDIA also released an open training dataset and reinforcement learning library, allowing developers to utilize Nemotron 3 in a variety of environments.

Unslose is a globally popular LLM fine-tuning framework that improves the performance of the HuggingFace Transformer library by up to 2.5x on NVIDIA GPUs.

There are three main types of fine-tuning methods.

Parameter-efficient fine-tuning (LoRA, QLoRA) is a method to improve performance quickly and cost-effectively by updating only part of the model.

Full fine-tuning is suitable for training a model to follow a specific format or style by updating all parameters.

Reinforcement learning is an advanced method that adjusts model behavior through feedback, and is used to improve accuracy in specialized fields such as law and medicine.

Unslose minimizes GPU memory usage to deliver optimized performance across a wide range of NVIDIA hardware, including RTX desktops and laptops, RTX PRO workstations, and DGX Spark.

NVIDIA supports local fine-tuning with DGX Spark based on the Grace Blackwell architecture.

DGX Spark provides up to 1 petaflop (FP4) of AI performance and 128GB of integrated memory, enabling it to reliably process large models with over 30 billion parameters.

Run fully fine-tuned and reinforcement learning-based workflows quickly, without waiting in the cloud. The advantage is that it can be controlled locally.

Additionally, high-resolution diffusion models can generate 1,000 images in just seconds, maintaining high throughput even in creative work or multimodal pipelines.

NVIDIA's announcement marks a significant turning point in realizing the potential of generative AI and agentic AI.

This enables the creation of a variety of customized AI assistants, including product support chatbots and personal schedule management secretaries, offering new opportunities for both businesses and individuals.

The combination of Nemotron 3, Unsloss, and DGX Spark will significantly increase the efficiency and accessibility of AI fine-tuning, making it easier for researchers and developer communities to realize AI innovation.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
명세환 기자
명세환 기자