This page was machine-translated and may differ from the original. View original

[2025 e4ds Tech Day] “Lightweight Technology is Key to the On-Device AI Era”

Google 우선 소스Published2025.10.20 10:50

Jo Seok-young, AI Manager at Nota, is giving a presentation on 'Hardware Cognitive AI Model Optimization Solutions and Success Stories for On-Device AI' at the '2025 e4ds Tech Day'.

Presenting the optimal level of lightweighting by considering the parallel processing characteristics of each hardware
Operator list pre-analysis, automatic conversion of replaceable operators, optimization

Lightweighting is not simply about making models smaller. Running AI in real-time on low-power devices requires complex technologies that reduce computational load and optimize for the hardware. Especially in on-device AI environments, lightweighting is a necessity, not an option, because AI must be executed on the device itself without relying on the cloud.

Jo Seok-young, AI Manager at Nota, gave a presentation on "Hardware Cognitive AI Model Optimization Solutions and Success Stories for On-Device AI" at the "2025 e4ds Tech Day" event held on September 9, highlighting the importance of lightweight technology for AI models in the era of on-device AI.

As AI technology advances rapidly, the size and complexity of models are increasing exponentially.

In particular, since the emergence of large language models such as GPT, AI models have been growing tenfold every two years, whereas hardware performance has only grown twofold every two years, following Moore's Law.

'AI model lightweighting' technology is attracting attention as a solution to bridge this gap.

Nota AI's lightweight technology consists of four main components.

The first is 'structural pruning'. Conventional unstructured pruning reduces computational load by replacing unimportant weights with zeros, but in actual hardware, the computational structure remains intact, so the speed improvement effect is limited.

On the other hand, Nota AI substantially reduces computational load and maximizes hardware performance through structural pruning that reduces the number of channels itself.

The second is hardware-specific optimization.

It is difficult to improve performance by simply making AI models smaller. To maximize the speed of emergence, the optimal level of lightweighting must be set by considering the parallel processing characteristics of each hardware.

Nota AI analyzes the stepped computational structures of various hardware, such as CPUs, GPUs, and DSPs, to achieve maximum performance improvement with minimal compression.

The third is operator conversion technology.

Complex AI models require advanced mathematical operators such as roots and logarithms, but hardware from many semiconductor startups cannot support them.

Nota AI analyzes the list of operators for each hardware in advance and optimizes the model by automatically converting unsupported operators into alternative methods.

There are also cases where this has made models that were not previously run on the MPU executable and dramatically improved the speed of emergence.

The fourth is post-correction technology after quantization.

Quantization is a technique that simplifies models to increase speed, but the decrease in accuracy is a problem. Nota AI minimizes accuracy loss due to quantization through proprietary post-correction technology and restores performance close to the original model.

In addition, it supports a mixed-precision method that quantizes only some layers, providing optimization tailored to customer needs.

This technological capability is leading to results in actual industrial settings.

Nota AI collaborates with global semiconductor companies such as ARM and Renesas, and has a track record of making models that previously could not run on MPUs runnable and improving speed by more than five times.

In addition, it has been recognized for its technological capabilities by attracting investment from major domestic companies such as Samsung Electronics, LG, Naver, and Kakao, as well as VCs such as SoftBank and Stonebridge.

Manager Cho Seok-young stated, “On-device AI offers various advantages, such as privacy protection, reduced latency, and minimized network dependency. However, lightweight technology must be in place to support its implementation,” adding, “This is because while cloud-based AI can utilize high-performance servers, on-device AI must operate on devices with limited memory and computing power.”

He also stated, “Ultimately, AI lightweighting technology is the key to realizing on-device AI,” adding, “These technologies, which reduce model size, optimize computations, and adapt to hardware, enhance the accessibility and efficiency of AI and enable practical applications across various industries.”
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
명세환 기자
명세환 기자