This page was machine-translated and may differ from the original. View original
Maintaining 98% of Original Performance Even After Removing 25% of Expert Modules from 31-Billion Parameter Model
Nota, an AI model optimization company, has developed technology that reduces GPU memory usage by nearly half while maintaining almost identical performance of multimodal AI models. The company announced that compression is possible without additional training, enabling improved efficiency in enterprise AI infrastructure operations.
Nota announced on the 1st that a paper containing lightweight techniques specialized in the Mixture of Experts (MoE) structure of multimodal AI models has been accepted to the main track of 'NeurIPS 2026', a world-renowned AI academic conference.
The accepted paper is titled 'Modality-Aware Expert Pruning for MoE-Based Multimodal Large Language Models'.
When existing expert pruning methods for language models were directly applied to multimodal models, expert modules necessary for image understanding were also removed, causing significant performance degradation in visual tasks such as chart analysis.
Nota separately evaluated expert importance for text and images respectively, using text-based importance criteria for selecting experts to retain and image-based importance criteria for selecting removal targets.
Instead of removing an equal number of experts at each layer, the company applied a method that prioritizes preserving experts in important layers by comparing the entire model.
According to performance evaluation results, when 25% of expert modules were removed from a multimodal model of approximately 31-billion parameters, 98% of the original performance was maintained, and when 50% were removed, 91% was maintained respectively.
GPU memory usage decreased from 58GB to 31GB, approximately 47% reduction, making it operable with a single GPU of 40GB memory.
The company explained that users need only set the removal ratio, and lightweight optimization can be completed with a single compression without additional training.
Nota plans to incorporate this technology into its AI optimization platform 'Netspresso' to expand support for multimodal and MoE models.
The company has previously applied expert pruning to super-large models such as 'Kimi K3' and 'Sola Open 2', and this year has had its research achievements recognized at major AI conferences including ICLR, EMNLP, and ICML, in addition to NeurIPS.
Chae Myung-soo, CEO of Nota, stated, "As the transition to multimodal AI accelerates, the burden of GPU memory is also increasing rapidly," and added, "This research is significant in that it can greatly reduce computing resources while maintaining model performance, thereby improving the efficiency of actual AI infrastructure operations."
To request a correction, reply or follow-up report on this article, see how to file a request. Previously published statements are collected in corrections & replies.













