
Limitations of 'Cat Body with Elephant Skin' Evaluation Method Revealed
A domestic research team has identified the limitations of a representative benchmark widely used to evaluate the visual recognition capabilities of artificial intelligence (AI) and proposed new evaluation standards.
The research team led by Professor Yoo Jae-joon from UNIST's Graduate School of Artificial Intelligence analyzed the structural problems of the existing AI visual recognition evaluation method 'Cue-conflict' and developed an improved new benchmark 'REFINED-BIAS,' announced on the 1st.
The research results were adopted as a spotlight paper at the European Conference on Computer Vision (ECCV 2026), selected among only approximately 1.3% of the 10,473 total submitted papers.
Cue-conflict is an AI evaluation method proposed in 2019 that utilizes data in which shape and texture are intentionally conflicted, such as an image of a cat shape with elephant skin texture. If the AI answers cat, it is evaluated as being dependent on shape; if it answers elephant, it is evaluated as being dependent on texture.
In the prior research, it was found that AI, unlike humans, relies more on texture than shape, which had a considerable impact on the learning direction of AI models.
Structural Problems of Existing Evaluation Methods Highlighted
The research team analyzed that cue-conflict has structural limitations that can distort the actual recognition characteristics of AI.
Notably, the existing method only evaluates the relative ratio of correct answers for shape and texture, failing to distinguish how much AI actually utilized the two types of information. Additionally, in some synthetic images, shape and texture information were not clearly separated, or biases existed where certain information was more easily recognized.
The scoring method was also limited. When the result that AI actually predicted with the highest confidence was excluded from evaluation targets, the result of the next rank was processed as if it were correct, creating a problem.
Measuring the Degree of Shape and Texture Utilization
The REFINED-BIAS proposed by the research team applied a method of scoring the degree of shape and texture utilization separately instead of the existing preference concept.
To accomplish this, 10 objects with distinct shapes and 10 objects with distinct textures were selected to construct a total of 6,000 evaluation images. Shape images had texture removed leaving only outlines, while texture images were processed so that the overall shape was not visible.
Experimental results showed that the more advanced AI models with superior performance actively utilized both shape and texture. Additionally, it was confirmed that recent transformer-based models and models trained on both images and text more effectively recognize shape information.
Professor Yoo Jae-joon stated, "Since benchmarks affect not only AI performance evaluation but also development direction, reliable evaluation standards are important," and added, "this research is expected to help more accurately analyze AI's visual recognition capabilities."
The research results were published under the title "On the Reliability of Cue Conflict and Beyond."













