KITECH Webinar_~7.22
Physical AI HBM Smart Factory SDV AIoT Power Semicon 특수 가스 정정·반론보도 모음 e4ds plus

NOTA Presents Two MoE Quantization Research Papers at ICML 2026 Workshop

Google 우선 소스 기사입력2026.06.11 10:32



DREAM-MoE and SRA-MoE Paper Accepted; Technology Presented to Improve Large-Scale AI Model Inference Efficiency
Nota, a company specializing in AI model lightweighting and optimization technology, will present two papers related to Mixture-of-Experts (MoE) model quantization at the 'Resource-Adaptive Foundation Model Inference (AdaptFM)' workshop at ICML 2026.

Nota plans to present its DREAM-MoE and SRA-MoE research at the AdaptFM workshop during ICML 2026, which will be held at COEX in Seoul from July 6 to 11. ICML is one of the leading international conferences in the field of machine learning, and AdaptFM is a workshop that covers inference, compression, and optimization techniques for efficiently running foundation models on limited computing resources.

MoE is an AI model architecture that selects and uses the necessary parts from among various expert models. While computational efficiency can be improved by not having to utilize the entire model every time, it requires an optimization method different from general models because it involves an expert selection process.

According to Nota, DREAM-MoE proposes a method to reduce variations in expert selection that can occur when quantizing a model by dividing it into multiple intervals. Quantization is a technique that reduces memory usage and computational burden by converting the numerical representation of an AI model to a lower precision.

SRA-MoE is a method that prioritizes the protection of inputs that have a greater impact on model results. Rather than treating all inputs equally, it is designed to ensure that expert selections for critical inputs do not vary significantly, focusing on maintaining model quality even with limited resources.

The company stated that the two studies demonstrated higher performance than existing MoE-specialized quantization techniques. This research can be utilized to mitigate quality degradation while reducing the memory and computational resources required to operate large-scale AI models.

Nota previously won both the track and overall titles at the NVIDIA Nemotron Hackathon with its data-driven MoE quantization technique. The company stated that it is also pursuing research on large-scale model optimization, including Solar MoE, within the Upstage Consortium's proprietary Foundation Model project.

Meanwhile, Nota plans to host 'Nota AI - Korea Efficient Days' at COEX in Seoul during ICML 2026 and introduce related research and application cases.