
▲ (From left) UNIST Professor Sung-Hwan Yoon, Researcher Jae-Jun Lee
The combination of different data mitigates the abruptness of error changes
The reason why multimodal artificial intelligence (AI) outperforms single-data-based AI has been mathematically explained. The principle by which the stability and generalization ability of a model improve when learning images, voice, and text together has been theoretically identified.
UNIST (Ulsan National Institute of Science and Technology) announced on the 26th that a research team led by Professor Sung-Hwan Yoon has identified the principles of performance improvement in multimodal AI from the perspective of 'Loss Landscape.' The research has been accepted for the International Conference on Machine Learning (ICML 2026).
Multimodal learning is a method of training AI by utilizing different forms of data together. Although cases of performance improvement have been reported previously, theoretical explanations linking it to the deep learning process have been limited.
The research team confirmed that when multimodal data is combined, the loss terrain becomes flatter during the model training process. Loss terrain is a concept representing the relationship between an AI model's error and internal parameters.
The wider and gentler the loss terrain, the less significantly performance is affected by new data or noise. The research team explained the reason for this flattening phenomenon occurring in multimodal learning as the 'Convolutional Smoothing Effect'.
The analysis suggests that combining different data mitigates the abruptness of error changes and enhances the model's robustness in responding to various situations.
Based on the theory, the research team proposed a new learning method called 'Distributional Multimodal Learning (DML).'
While existing methods train data such as images and text using fixed pairs, DML is a method that randomly reconstructs combinations of data that share the same meaning.
They explained that this allows for increased diversity in training data and enhanced lossy terrain flattening effects. Experiments with various multimodal datasets also showed improvements in classification accuracy and search performance compared to existing methods.
The research team stated, “We theoretically explained why multimodal AI generalizes more robustly and presented a new learning method.”
This research is presented as a foundational technology that can be utilized in various fields, including autonomous driving, medical AI, robotics, and foundation models.