KAIST develops AI semiconductor technology for deep reinforcement learning processing
Expected to be used in robot control, autonomous drones, games, etc. Professor Yoo Hoi-Jun's team at the Korea Advanced Institute of Science and Technology (KAIST) announced on the 16th that they have developed a semiconductor technology that can process deep reinforcement learning (DRL), which was used in Google's AI program, 'AlphaGo', with high performance and power efficiency. This research was introduced as a highlight paper at the 'IEEE VLSI Symposia' held from June 14 to 19.

▲ KAIST develops AI semiconductor technology OmniDRL [Capture = SSL KAIST]
Unlike 'supervised learning', where AI learns by using pairs of data and correct answers created by humans, DRL is a method where AI derives the optimal answer on its own by using the experience gained through trial and error in a given environment, and then humans provide feedback on the result.
It has the characteristic of using multiple neural networks simultaneously to quickly find the optimal answer in a situation where the correct answer is not given. However, since the neural networks are complexly intertwined and a large amount of data must be processed, it could only be implemented by utilizing multiple high-performance computers with large amounts of memory in parallel in the past. Therefore, laptops, smartphones, etc. could not implement this.
The research team developed 'OmniDRL', an AI semiconductor technology that has superior performance and 2.4 times higher power efficiency than existing technologies, enabling DRL on mobile devices.
Specifically, it used technology that increases the compression ratio of deep neural network data and reduces the amount of unnecessary or duplicated data, technology that enables calculations in a compressed state of data, unlike before, and SRAM-based processing-in-memory (PIM) semiconductor technology that integrates calculation and storage functions. In particular, existing PIM semiconductors could only perform operations on integers, but through this research, a technology was developed that enables decimal-based operations.
As a result of applying OmniDRL to the 'Humanoid Robot Adaptive Walking System', it was confirmed that adaptive walking was possible at a speed more than 7 times faster than when OmniDRL was not connected.
Professor Yoo Hoe-jun said, “This research enabled inference and learning while maintaining a deep neural network in a compressed state with a single semiconductor,” and explained, “It is significant in that we developed AI semiconductor technology that enables decimal point operations that were previously thought to be impossible.”
Song Kyung-hee, AI-based policy officer at the Ministry of Science and ICT, said, “This study is significant in that it has resulted in international recognition of domestic research results in the field of AI semiconductors,” and added, “The Ministry of Science and ICT will continue to support the 1 trillion won AI semiconductor R&D project that began last year, and will fully promote the 400 billion won PIM semiconductor R&D project starting next year.”
Meanwhile, this study was conducted with the support of the Ministry of Science and ICT's 'Innovation Growth Linked Intelligent Semiconductor Leading Technology Development' project, which has a total budget of 1.8 billion won from 2019 to 2021.