마이크로칩 8월
This page was machine-translated and may differ from the original. View original

Nota, 2 MoE Quantization Papers Accepted at EMNLP 2026

Google 우선 소스Published2026.08.25 09:49


 
Qwen 3.8 Max GPU Requirements Reduced from 24 to 4 Units, LLM and Data Center Optimization Technology Application Expanding
 
Nota AI, an AI model lightweight and optimization company, has had 2 papers accepted at EMNLP 2026, a major international conference in natural language processing (NLP), continuing to deliver LLM optimization research results. The company is expanding the scope of technology application to the data center AI infrastructure sector, such as reducing the required NVIDIA B300 GPUs for running Qwen 3.8 Max from 24 to 4 units.
 
Nota announced on the 25th that papers on large-scale AI model optimization have been accepted at both the Main Conference and Findings of EMNLP 2026, with one paper in each category.

This year's conference received approximately 18,000 submissions, with a Main Conference acceptance rate of 15.4%.

EMNLP is a leading academic conference in the NLP field where companies such as Google, Meta, and OpenAI present their latest research.
 
Both accepted papers address the performance degradation issue that occurs in the quantization process of Mixture of Experts (MoE) models.

MoE performs computations using only some experts depending on the input, but the entire model must be loaded into memory, requiring large-scale GPUs for actual operation.

Reducing this through quantization can change expert selection, potentially degrading answer quality.
 
The Main Conference accepted paper presented the 'MENDS-MoE' methodology, which reflects the impact of quantization on next-step expert selection.

According to Nota, experiments with 4-bit and 3-bit quantization on 3 types of MoE models showed higher average accuracy compared to competing technologies.

The Findings accepted paper proposed the 'OPERA' methodology, which focuses optimization only on expert selection changes that actually affect responses.
 
Prior to this paper acceptance, Nota ranked 3rd among approximately 40 teams worldwide at the ICML 2026 'AdaptFM' workshop 'Efficient Qwen Competition'.

The company demonstrated running Qwen 3.5-4B on a single NVIDIA A10G GPU with inference speed 6.978 times faster on average compared to existing solutions.

Two MoE quantization papers were also accepted at the same workshop.
 
The company also disclosed actual model application results.

△Qwen 3.8 Max (over 1 trillion parameters) B300 GPU from 24 to 4 units △Moonshot AI 'Kimi K3' B300 GPU from 8 to maximum 4 units △Upstage 'Solar Open 2' H100 GPU from 8 to 2 units can each be reduced, Nota explained.

To date, Nota has published approximately 50 research results across major domestic and international conferences and journals, from core technologies such as pruning and quantization to vision language models (VLM) and LLMs.
 
Nota CEO Chae Myung-soo stated, "As the competitive axis in the AI industry shifts beyond model performance to operational efficiency, this achievement demonstrates that we have secured the technical foundation to improve the efficiency of large language models and data center AI infrastructure," and added, "We will continue to expand the scope of application of core optimization technologies that enhance the efficiency of AI infrastructure adoption."
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
배종인 기자
배종인 기자