This page was machine-translated and may differ from the original. View original

▲ Speech recognition architecture diagram of the newly developed 'AIMZformer' (Image: Mediazen)
Significant improvement in commercial voice recognition processing speed expected
Significant improvement in commercial voice recognition processing speed expected
To provide a latency-free speech recognition solution, a speech recognition system with a convolutional network structure and 40% improved processing speed was developed.
Mediazen announced on the 26th that it has developed an augmented Transformer-based speech recognition system with a new convolutional network structure that benchmarks 'Conformer,' a representative E2E speech recognition system developed by Google, and can improve processing speed by about 40% while maintaining the performance of the existing Conformer.
This technology development was carried out through the Korea Electronics and Telecommunications Research Institute (ETRI)’s field support program for research personnel, with the participation of speech recognition experts including ETRI’s Principal Researcher Lee Seong-ju and Mediazen’s AIMZ Research Director Yoon Jong-seong.
As a result of speech recognition experiments using the LJSpeech dataset, Google’s conformer showed performance of CER 4.8% and WER 19.6%, and the tentatively named ‘AIMZformer’ (Mediazen speech recognition system) showed performance of CER 4.8% and WER 19.2%, respectively.
Based on this, it can be seen that speech recognition performance at the level of Google Conformer is maintained, and the processing speed has been significantly improved to 80ms compared to 40ms with Conformer subsampling. This saves about 40% of study time.

▲ Block diagram of newly developed speech recognition convolution (Image: Mediazen)
For reference, the baseline Transformer-based speech recognition system showed recognition performance of CER 6.9% and WER 23.0%. In this experiment, to evaluate the performance of the pure neural network, training and evaluation were conducted using only the categorical cross-entropy of the output nodes without performing backend processing such as beam search. In addition, it is known that the recognition difficulty is high because alphabet-based characters are used as the units required for voice recognition.
While Google's conformer focuses on encoder performance, the convolutional structure developed by Mediazen's AIMZ Lab focuses on versatility, offering the advantage of improving the performance of both encoders and decoders.
In particular, to secure competitiveness in embedded solution development, Mediazen is pursuing the supply of high-speed engines that can be installed in non-networked electronic devices, such as AI robots and small electronic devices.
Yoon Jong-sung, Director of Mediazen AIMZ Research Institute, stated, “With the development of this new technology, we have secured proprietary conformer technology that significantly increases processing speed while maintaining voice recognition performance, thereby greatly improving user satisfaction for those dissatisfied with response speeds.” He added, “In the future, speed improvements will be achieved across all businesses where voice recognition technology is utilized.”
Meanwhile, it has been reported that Mediazen’s AIMZ Lab is not only developing various technologies focused on speech recognition speed but also possesses the technology to perform multilingual speech recognition with a single model, and is simultaneously developing a new Large-Length Model (LLM) to achieve international technological competitiveness on par with global platform companies.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.














