71% Reduction in Model Size While Maintaining 99.2% Accuracy, Confirming the Potential of Sovereign AI Infrastructure
The possibility that a combination of domestic NPUs, large-scale AI models, and AI optimization technology can achieve actual service-level performance has been confirmed for the first time in Korea. Nota (CEO Chae Myeong-su), a company specializing in AI model lightweighting and optimization, has succeeded in optimizing LG AI Research Institute's K-EXAONE 236B on Furiosa AI's data center NPU, demonstrating the possibility of operating high-performance large language models (LLM) within the domestic AI ecosystem.
Nota announced on the 30th that it has optimized the K-ExaOne 236B, with a scale of approximately 236 billion parameters, for the FuriosaAI NPU environment.
Instead of readjusting the entire model, we chose an approach that precisely analyzes sections where performance degradation may occur and selectively applies optimization only to the necessary parts.
As a result, the company explained that it succeeded in reducing the model size by about 71% to lower the memory burden while maintaining the original level of accuracy in major benchmarks.
K-ExaOne 236B is a large AI model that adopts a Mixture of Experts (MoE) architecture that selectively utilizes multiple expert models.
While the MoE structure is advantageous for increasing model efficiency, it requires sophisticated techniques to ensure that each expert model operates stably during the optimization process.
In particular, minute errors occurring during the quantization process can affect the accuracy of the final answer if they accumulate over a long inference process.
Nota overcame these structural difficulties and scored 79.80 points in Scientific Reasoning (GPQA), 68.98 points in Instructional Understanding (IFBench), and 88.57 points in Math Problem Solving (AIME25).
The performance of the original model was 79.1, 67.3, and 92.8 points, respectively, and it was found that it maintained about 99.2% accuracy compared to the original based on the simple average of the three items.
Amid the emergence of the so-called 'Sovereign AI' trend in the global AI industry, which seeks to secure domestic AI models and computing infrastructure, this achievement is It was presented as an example demonstrating the potential for linking domestic AI semiconductors, domestic AI models, and optimization software.
"It is important that models, semiconductors, and optimization software are connected into a single actionable AI infrastructure," said Chae Myeong-su, CEO of Nota. "This achievement is a case that confirms the feasibility of operating large-scale AI models in practice by combining FuriosaAI's NPU, LG's K-ExaOne, and Nota's optimization technology."