Techday
Physical AI HBM Smart Factory SDV AIoT Power Semicon 특수 가스 정정·반론보도 모음 e4ds plus

UNIST Develops AI Technology to Find Objects in 3D Space Using Natural Language

Google 우선 소스 기사입력2026.06.08 15:01



LightSplat reduces recognition preparation time to 5 seconds and memory usage to 1/64th of the previous level.
A research team at UNIST has developed AI technology that locates specific objects within a three-dimensional space based on natural language input from users. The method identifies the location and area of an object in a 3D reconstructed space when the target is specified using concrete expressions, such as "white sofa" or "egg on top of ramen."

UNIST announced on the 8th that a team led by Professor Kyung-Don Joo of the Graduate School of Artificial Intelligence has developed 'LightSplat,' an open vocabulary-based 3D spatial recognition technology. The results of this research were accepted at CVPR 2026, an international conference in the field of computer vision held in Denver, USA, for five days starting from the 3rd.

LightSplat is a technology designed to rapidly perform natural language-based object recognition in fields dealing with three-dimensional space, such as robotics, augmented reality, and digital twins. Existing 3D spatial recognition technologies operated based on predetermined object categories or required storing a large amount of semantic information for each point particle in 3D space, resulting in high processing time and memory burden.

The research team applied a method of attaching short 2-byte indices to point particles, or Gaussians, in 3D space reconstructed from camera images, instead of directly storing long linguistic feature values. The actual semantic information is stored in a separate table and retrieved via the index when needed.

Through this, the time required to make a 3D space searchable was reduced to about 5 seconds. The research team explained that this is approximately 50 to 400 times faster than existing state-of-the-art technology, and memory usage has decreased to one-sixty-fourth of the level.

In the performance evaluation, experiments were conducted to distinguish between small and distant objects using the LERF-OVS and DL3DV-OVS datasets. Results were confirmed regarding the identification of objects of different sizes and arrangements, such as an egg on top of ramen, tea inside a glass, a distant car, and office furniture.

In the ScanNet-based 3D semantic segmentation experiment, an mIoU of 37.11 was recorded across 19 classification criteria. mIoU is an indicator of how much the object region predicted by the AI overlaps with the actual correct answer region.

"To actually use Open Vocabulary 3D object recognition technology, we must secure not only accuracy but also speed and memory efficiency," said researcher Jaehoon Bang, the first author.

Professor Joo Kyung-don stated that this technology can be utilized in robots that understand natural language instructions, AR/VR content creation, and digital twin-based spatial management. The research was conducted with support from the Ministry of Science and ICT and the Korea Institute of Information and Communication Technology Planning and Evaluation.