Techday
This page was machine-translated and may differ from the original. View original

NOTA Takes 3rd Place at ICML 2026 Challenge… Qwen Inference Optimization Technology Validated

Google 우선 소스Published2026.07.13 09:00

6.978x improvement in inference speed in a single A10G GPU environment
Nota AI, a company specializing in AI model lightweighting and optimization, took third place in the 'Efficient Qwen Competition' held at the global machine learning conference ICML 2026. In this competition, in which over 40 teams from around the world competed, Nota increased inference speed by an average of 6.978 times with a technology that combines quantization and speculative decoding, and two related papers were also accepted at the ICML AdaptFM workshop.

This competition is a challenge under the ICML 2026 AdaptFM (Resource-Adaptive Foundation Model Inference) workshop, with researchers from global companies such as Amazon and Meta participating as organizing committee members.

The participating teams performed the task of maximizing inference speed while maintaining the answer quality of the open-source LLM Qwen3.5-4B in a single NVIDIA A10G GPU environment.

Nota reduced the model's memory usage and computational load through quantization, and then minimized the degradation of accuracy through subsequent training.

They stated that they combined speculative decoding, where a draft model rapidly generates answer candidates and the main model validates them, and further reduced unnecessary computations by applying the Sliding-window Attention technique.

It is reported that many of the top teams in the tournament have adopted similar strategies.

Kim Tae-ho, CTO and co-founder of Nota, stated, “This is a case where our inference optimization technology has been validated using Qwen, a representative open-source model utilized in the global AI ecosystem,” adding, “We will expand the application of optimization technology to on-device and edge AI environments.”

Two papers accepted at the AdaptFM workshop deal with MoE (Mixture of Experts) structure LLM quantization.

An optimization methodology that reduces performance degradation with limited memory and computational resources by leveraging the structural characteristics of MoE.This is the content of the proposal.

Nota previously won both the track and the overall title at the NVIDIA Nemotron hackathon using MoE quantization technology.

During ICML 2026, Nota held 'Nota AI - Korea Efficient Days' near COEX in Samseong-dong, Seoul.

Representatives, researchers, and engineers from global companies such as OpenAI, Google, and Qualcomm attended the event to discuss research trends and industrial application possibilities for Efficient AI.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
배종인 기자
배종인 기자