This page was machine-translated and may differ from the original. View original

Nota Unveils Upstage's "Solar Open 100B" Quantization Model, Increasing Weight Memory from 191.2GB to 51.9GB

Google 우선 소스Published2026.03.05 10:31


Results of MoE-specific quantization revealed amid the race to reduce the weight of large LLMs.

As demand for running Large Language Models (LLMs) on field devices or in limited GPU environments increases, competition for 'quantization' technology that reduces memory usage without significantly compromising model accuracy is intensifying.

Nota announced on March 5 that it had applied its quantization techniques to Upstage's 'Solar-Open-100B' to reduce memory usage.

According to the model card released by Nota on Hugging Face, the model weight memory footprint of Solar-Open-100B with 'Nota MoE Quantization' applied was presented as 51.9GB, down from 191.2GB. The same page explained that this approach is designed to mitigate **distortion** that can occur in the process of mixing experts in MoE (Mixture of Experts) structure-based LLM.

**Perplexity (PPL)** was also presented as a performance indicator. The table on the model card shows the PPL based on WikiText-2 as 6.06 for the original Solar-Open-100B and 6.81 for the Nota technique-applied model. PPL is used to indicate that the lower the value, the more stable the language model's predictions are.

This disclosure is noteworthy because, amid the proliferation of MoE-affiliated LLMs, it goes beyond a "universal reduction of the entire model" approach and instead presents a quantization approach that considers structural characteristics as an indicator of actual model card performance. However, some media reports have also highlighted the Ministry of Science and ICT's "Independent AI Foundation Model" project and whether a patent has been filed, and the extent of verification based on articles and publicly available data may vary.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
배종인 기자
배종인 기자