Techday
자동차 / 모빌리티 산업자동화 / 로봇 가전 / 스마트홈 의료기기 / 바이오 통신 / 네트워크 에너지 / 전력 / ESS 디스플레이 / 광학 반도체 / 부품소자 국방 / 항공우주 철도 / 중장비 / 중공업 교육 / 리서치 / 공공 스마트시티 / 건물자동화 해양 / 수처리 / 선박 소비자 전자 / 웨어러블 기타

UNIST Identifies Principles of Multimodal AI Performance Improvement… Paper Accepted for ICML 2026

Google 우선 소스 기사입력2026.06.29 14:16


▲ (From left) UNIST Professor Sung-Hwan Yoon, Researcher Jae-Jun Lee

The combination of different data mitigates the abruptness of error changes

The reason why multimodal artificial intelligence (AI) outperforms single-data-based AI has been mathematically explained. The principle by which the stability and generalization ability of a model improve when learning images, voice, and text together has been theoretically identified.

UNIST (Ulsan National Institute of Science and Technology) announced on the 26th that a research team led by Professor Sung-Hwan Yoon has identified the principles of performance improvement in multimodal AI from the perspective of 'Loss Landscape.' The research has been accepted for the International Conference on Machine Learning (ICML 2026).

Multimodal learning is a method of training AI by utilizing different forms of data together. Although cases of performance improvement have been reported previously, theoretical explanations linking it to the deep learning process have been limited.

The research team confirmed that when multimodal data is combined, the loss terrain becomes flatter during the model training process. Loss terrain is a concept representing the relationship between an AI model's error and internal parameters.

The wider and gentler the loss terrain, the less significantly performance is affected by new data or noise. The research team explained the reason for this flattening phenomenon occurring in multimodal learning as the 'Convolutional Smoothing Effect'.

The analysis suggests that combining different data mitigates the abruptness of error changes and enhances the model's robustness in responding to various situations.

Based on the theory, the research team proposed a new learning method called 'Distributional Multimodal Learning (DML).'

While existing methods train data such as images and text using fixed pairs, DML is a method that randomly reconstructs combinations of data that share the same meaning.

They explained that this allows for increased diversity in training data and enhanced lossy terrain flattening effects. Experiments with various multimodal datasets also showed improvements in classification accuracy and search performance compared to existing methods.

The research team stated, “We theoretically explained why multimodal AI generalizes more robustly and presented a new learning method.”

This research is presented as a foundational technology that can be utilized in various fields, including autonomous driving, medical AI, robotics, and foundation models.