마이크로칩 8월
This page was machine-translated and may differ from the original. View original

[Serial Project] Physical AI: Redefining the Future of Industry ⑥, "Multimodal AI: The 'Brain' and 'Language' of the Physical AI Era"

Google 우선 소스Published2025.12.09 11:14
"Multimodal AI: The Brain and Language of Physical AI"
A core technology for general intelligence that simultaneously understands robot language, vision, and behavior.
Global companies like Google, Nvidia, and Figure AI are fiercely competing in technology.


[Editor's Note] While conventional AI focuses on inference and generation within digital data, Physical AI directly acts and reacts in the real world through sensors, edge computing, robots, and control systems. The implementation of Physical AI can significantly advance industrial innovation and automation because it directly acts and solves problems in the real world, and it directly interacts with the real world. Accordingly, global companies such as Nvidia, Tesla, and Google are making massive investments in Physical AI, and the related market is expected to grow explosively. To implement this Physical AI, recognition technologies such as sensors, as well as edge computing and embedded systems for real-time data processing, local computing, robotics, and control technologies are essential. Accordingly, e4ds News has prepared a series of articles to examine the core technologies and implementation strategies of physical AI, including the concept, market outlook, related technologies, and actual cases.

▲Photo: pixabay.com


Artificial intelligence (AI) is moving beyond simply analyzing data and processing language to enter the era of physical AI, where it directly interacts with the physical world.

This trend, in which robots perceive the world through sensors, move with actuators, and autonomously perform tasks in real-world environments, symbolizes the transition from “thinking AI” to “acting AI.”

The most crucial technology that emerged in this process is multi-modal AI.

Multimodal AI refers to artificial intelligence that simultaneously learns and understands different forms of information, such as text, images, videos, and behavioral data.

While existing language models only processed text, multimodal AI integrates vision, language, and behavior to enable robots to understand and execute commands in a more human-like manner.

For example, it can process not only simple instructions like “Pick me an apple,” but also abstract, semantic commands like “I’m hungry, find me a healthy snack.”

This is a key to enabling robots to go beyond simple repetitive tasks and acquire generalized intelligence that can respond to unpredictable situations.

It is clear why multimodal AI is important in implementing physical AI.
/> Robots simultaneously receive various sensory data from the real world.

The visual information provided by the camera, the spatial data generated by LiDAR, the pressure and texture detected by the tactile sensor, and the verbal instructions given by the user must all be combined.

Multimodal AI connects these disparate data sets into a single, unified intelligence, enabling robots to behave naturally in real-world environments.

Ultimately, multimodal AI serves as the 'brain' of physical AI.

Global companies are already engaged in fierce competition centered around multimodal AI.

Google DeepMind has unveiled a vision-language-action (VLA) model that combines web-scale data and robotics data through 'RT-2 (Robotic Transformer 2)'.

RT-2 demonstrated "creative abilities," such as recognizing new objects not included in existing training data and interpreting abstract commands, opening up new possibilities for robotic intelligence.

NVIDIA is developing a universal humanoid foundation model called 'GR00T' based on its hardware and simulation ecosystem.

GR00T learns from massive amounts of synthetic data generated from NVIDIA's simulation platform Isaac Sim, and acquires various behaviors through text and image prompts.

This demonstrates that multimodal AI is not simply limited to software, but is at the center of a platform competition encompassing hardware, simulation, and data pipelines.

Startup Figure AI is also developing a general-purpose humanoid robot using multimodal AI.

theseOnce, in collaboration with OpenAI, a robot equipped with a vision-language model (VLM) as its brain was presented that could converse with humans in natural language, recognize the surroundings, and act accordingly.

Although the collaboration ended shortly, the combination of "the best AI brain + agile hardware" is considered a symbolic example of the future of physical AI.

▲Comparison of Motimodal AI Strategies by Company


The future development of multimodal AI is expected to determine the success or failure of the commercialization of physical AI.

First, generating synthetic data through simulation will become more important.

Collecting massive amounts of behavioral data in the real world is significantly limited by cost and safety concerns.

Therefore, it is likely that mass production and learning of multimodal data in a high-fidelity digital twin environment will become mainstream.

Second, as cross-embodiment technology advances, a single multimodal AI model will be able to be applied to various robot platforms.

This greatly increases the versatility of the robotics industry, enabling a wide range of applications from smart factories to home service robots.

Third, the evolution of communications infrastructure will also accelerate the spread of multimodal AI.

5G and 6G networks support ultra-low latency connections between robots and cloud brains, providing the foundation for multimodal AI to operate in real time.

Ultimately, multimodal AI is the 'brain' of the physical AI era. It's 'language'.

Multimodal AI is essential for robots to see, hear, understand, and act in the same way as humans.

The strategies of global companies are expanding beyond simple technological competition to platform competition that integrates hardware, software, data, and communications.

The future of physical AI depends on how quickly and sophisticatedly multimodal AI can integrate with the real world.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
배종인 기자
배종인 기자