마이크로칩 8월
This page was machine-translated and may differ from the original. View original

ETRI Unveils "Covert," a Korean Language Model for AI Applications

Google 우선 소스Published2019.06.11 09:50
| 4.5% better performance on average compared to Google's Korean language model
| AI can be used in Korean language processing fields such as assistants and Q&A.
| Available in both PyTorch and TensorFlow environments


The Exobrain project, an innovation growth engine project promoted by the Ministry of Science and ICT and the Institute of Information and Communications Technology Planning and Evaluation (IITP), has unveiled a cutting-edge Korean language model. This is expected to further advance the development of Korean-based AI services, including AI assistants, AI question-and-answer, and intelligent search.
ETRI Exobrain logo

On the 10th, the Electronics and Telecommunications Research Institute (ETRI) released the cutting-edge Korean language model 'KorBERT' through its website .

A language model is a corpus that represents language as numbers for natural language processing deep learning and then collects the probability distribution of words that appear according to learning.

The researchers released two models: one based on Google's language representation method, with more Korean data, and another that incorporates the agglutinative nature of Korean.

This technology was incorporated into the beta version of Hancom Office Knowledge Search in March of this year. In the second half of the year, we plan to additionally release a 'Law Field Question-Answering API (Application Programming Interface)' utilizing ETRI's language model, and also launch 'Similar Patent Intelligent Analysis Technology.'
ETRI Senior Researcher Lim Jun-ho explains how Covert works.

Developing deep learning technologies for language processing requires representing phrases in text as numbers. To this end, organizations developing language-based services have primarily used Google's multilingual language model, BERT.

Bert divides phrases within sentences into individual letters, then recognizes frequently occurring letters as words. This method, first released in November of last year, garnered significant attention for its significant performance improvements across 11 areas of language processing.

Google developed a Korean language model using data from over 400,000 Wikipedia articles. ETRI researchers then added 23 gigabytes (GB) of newspaper articles and encyclopedias from the past decade, training a total of 4.5 billion morphemes to develop a language model based on more Korean data than Google.

Simply increasing the amount of input data has limitations in improving language model development. Unlike other languages, Korean is an agglutinative language, with particles attached to roots. The research team created a language model that reflected the characteristics of the Korean language to the greatest extent possible, taking into account even the morpheme, the smallest unit of meaning in the Korean language.
Comparison table of algorithms between Covert and Google Language Model
Comparison table of algorithms between Covert and Google Language Model

The research team explained that the features that differentiate this Korean-optimized language model from Google are ▲a language model that analyzes morphemes in the preprocessing process ▲learning parameters optimized for Korean ▲a vast data base.

The developed language model outperformed the Korean model distributed by Google by an average of 4.5% across five performance criteria. In terms of passage ranking, it recorded a high figure of 7.4%.

The research team's language model is expected to be widely utilized by developers at universities, companies, and institutions for purposes such as deep learning research and education, as it can improve service performance and competitiveness.

The developed language model can be used in both PyTorch and Tensorflow environments, which are representative deep learning frameworks, and can be easily found in the public AI open API/data service portal.

ETRI's Dr. Kim Hyun-ki, who is in charge of the Exobrain project, said, "We expect that various Korean deep learning technologies, such as Korean analysis, knowledge inference, and question-answering, will be advanced through a language model optimized for the Korean language."

Kim Ji-won, head of the Ministry of Science and ICT's Artificial Intelligence Policy Team, also said, "We will strive to promote open innovation by disclosing high-quality artificial intelligence software APIs and data developed through government R&D through the AI Hub."

The current method used by Google and its researchers to develop language models cannot process documents containing more than 512 words at a time. In the future, the research team plans to develop a model that can process more language data at once and utilizes more advanced verification methods.
ETRI's Exobrain Technology Wins the 2016 EBS Scholarship Quiz

The 'Exobrain Project', which formed the basis of this research and development, has achieved successes such as winning the 2016 EBS Scholarship Quiz, 39 cases of technology transfer and commercialization, 44 cases of domestic and international standardization, and 70 patent applications.

Meanwhile, ETRI has been releasing open APIs and machine learning data for language intelligence technology since 2017. To date, over 13 million applications have been made, with developers from industries (42%), universities (34%), individuals (20%), and others (4%) using the technology. Furthermore, ETRI is developing AI-focused services for banks and local governments, promoting the industrialization of AI in Korea.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
이수민 기자