Tektronix TIF 2026
This page was machine-translated and may differ from the original. View original

ETRI: Input a sentence and get an image in just 2 seconds.

Google 우선 소스Published2024.01.29 11:56

▲ETRI researchers demonstrate the KOALA model, which creates images by inputting sentences.

Ultra-fast generative visual intelligence model unveiled, five times faster than Dali.
Three image generation models and two interactive visual language models released.

A domestic research team is releasing to the public a technology that combines generative artificial intelligence and visual intelligence technology to create images in just two seconds when a sentence is entered. Research on ultra-fast generative visual intelligence is expected to gain momentum.

The Electronics and Telecommunications Research Institute (ETRI) announced on the 26th that it will release a total of five models to the public, including three models of 'KOALA' that are five times faster than existing models in creating images by inputting sentences, and two models of 'Ko-LLaVA', an interactive visual language model that can load images or videos and answer questions.

First, the 'KOALA' model dramatically reduced the 2.56B (2.5 billion) parameters of the open source model to 700M (700 million) by applying the knowledge distillation technique.

If the number of parameters is large, the amount of calculation increases, which takes a long time and increases the service operation cost.

The researchers reduced the model size by one-third and improved the high-resolution image processing speed by about two times compared to the previous model and five times compared to DALL-E 3.

ETRI announced that it has reduced the model generation speed to around 2 seconds and significantly reduced the model size, so that it can be run on low-cost graphics processing units (GPUs) with only 8GB of memory amidst the recent domestic and international competition to create images from sentences (text).

The three parameter-specific 'KOALA' models developed by ETRI were unveiled in the HuggingFace environment.

In fact, when the researchers input the sentence “Photo of an astronaut reading a book on Mars under the moon,” ETRI’s Koala 700M (700 million pixels) created the image in just 1.6 seconds.

Calo (Kakao Brain) took 3.8 seconds, Dali 2 (Open AI) took 12.3 seconds, and Dali 3 (Open AI) took 13.7 seconds.

ETRI has two existing open source S/W stable diffusion models, BK-SDM and Kal released by the company.We have created and released a website where you can directly compare and experience a total of 9 models, including 4 types: Karlo, DALL-E 2, and DALL-E 3, as well as a website that provides models.

The research team also unveiled the 'Ko-LLaVA' model, a conversational visual language model that adds visual intelligence technology to conversational AI such as ChatGPT and can load images or videos and ask and answer questions about them in Korean.

The 'Lava' model was developed through international joint research between researchers at the University of Wisconsin-Madison and ETRI.

It was presented at NeurIPS'23, the top conference in the field of artificial intelligence, and used open source LLaVA, which has image interpretation capabilities at the level of GPT-4.

The research team conducted extended research based on the Lava model, which is emerging as an alternative to multimodal models including images, to better understand Korean and perform video interpretation that did not exist before.

In addition, the company also pre-released its self-developed Korean-based small language understanding-generation model (KEByT5).

The released models (330M (Small), 580M (Base), 1.23B (Large) class) applied token-free technology that can process new words and unlearned words.

Learning speed was enhanced by more than 2.7 times, and inference by more than 1.4 times.

The research team predicts that the current generative artificial intelligence market is gradually shifting from sentence-based generative models to multimodal generative models, and that smaller and more efficient models will emerge in the competition for model size.

The reason ETRI is releasing this model is that if the model is large, thousands of servers are needed, but by reducing the model, small and medium-sized enterprises can use it, thereby creating a related market ecosystem.The intention is to do so.

In the future, the research team predicts that there will be a high demand for a Korean cross-modal model that adds visual intelligence technology to the representative open language model of generative AI.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
배종인 기자
배종인 기자