나는엔지니어다
This page was machine-translated and may differ from the original. View original

UNIST Identifies Flaws in AI Visual Recognition Evaluation Standards

Google 우선 소스Published2026.10.01 09:49


Limitations of 'Cat Body with Elephant Skin' Evaluation Method Revealed


A domestic research team has identified the limitations of a representative benchmark widely used to evaluate the visual recognition capabilities of artificial intelligence (AI) and proposed new evaluation standards.


The research team led by Professor Yoo Jae-joon from UNIST's Graduate School of Artificial Intelligence analyzed the structural problems of the existing AI visual recognition evaluation method 'Cue-conflict' and developed an improved new benchmark 'REFINED-BIAS,' announced on the 1st.


The research results were adopted as a spotlight paper at the European Conference on Computer Vision (ECCV 2026), selected among only approximately 1.3% of the 10,473 total submitted papers.


Cue-conflict is an AI evaluation method proposed in 2019 that utilizes data in which shape and texture are intentionally conflicted, such as an image of a cat shape with elephant skin texture. If the AI answers cat, it is evaluated as being dependent on shape; if it answers elephant, it is evaluated as being dependent on texture.


In the prior research, it was found that AI, unlike humans, relies more on texture than shape, which had a considerable impact on the learning direction of AI models.


Structural Problems of Existing Evaluation Methods Highlighted


The research team analyzed that cue-conflict has structural limitations that can distort the actual recognition characteristics of AI.


Notably, the existing method only evaluates the relative ratio of correct answers for shape and texture, failing to distinguish how much AI actually utilized the two types of information. Additionally, in some synthetic images, shape and texture information were not clearly separated, or biases existed where certain information was more easily recognized.


The scoring method was also limited. When the result that AI actually predicted with the highest confidence was excluded from evaluation targets, the result of the next rank was processed as if it were correct, creating a problem.


Measuring the Degree of Shape and Texture Utilization


The REFINED-BIAS proposed by the research team applied a method of scoring the degree of shape and texture utilization separately instead of the existing preference concept.


To accomplish this, 10 objects with distinct shapes and 10 objects with distinct textures were selected to construct a total of 6,000 evaluation images. Shape images had texture removed leaving only outlines, while texture images were processed so that the overall shape was not visible.


Experimental results showed that the more advanced AI models with superior performance actively utilized both shape and texture. Additionally, it was confirmed that recent transformer-based models and models trained on both images and text more effectively recognize shape information.


Professor Yoo Jae-joon stated, "Since benchmarks affect not only AI performance evaluation but also development direction, reliable evaluation standards are important," and added, "this research is expected to help more accurately analyze AI's visual recognition capabilities."


The research results were published under the title "On the Reliability of Cue Conflict and Beyond."

To request a correction, reply or follow-up report on this article, see how to file a request. Previously published statements are collected in corrections & replies.
배종인 기자
배종인 Reporter

Comments