Robot Brain Race Enters Deep Water: Why Physical World Requires Designing Models from Scratch
At the 2025 WAIC, embodied AI took center stage with over 200 companies exhibiting, shifting focus from performance robots to practical applications like handling and inspection. Ant Lingbo's chief scientist, Shen Yujun, argues that adapting digital world video models to robots via fine-tuning is a path-dependent shortcut, not an ultimate solution, advocating for an 'embodied-native' approach with models designed from the ground up for physical reality. The company unveiled its full-stack Brain 2.0 with six models, including the industry-first embodied-native world-action model, and announced a financing round aiming to raise 1.5 billion yuan.
In an interview at WAIC, Ant Lingbo's chief scientist Shen Yujun discussed the challenges of embodied intelligence, emphasizing that data standards have not converged because technical routes are still undecided, with no consensus on modalities, precision, or quality. He noted that real data, especially teleoperation data from robots and first-person data from humans, is more critical than simulation data for general-purpose robots, as simulations struggle to replicate human subconscious actions. Lingbo's strategy involves developing multiple models (like LingBot-Vision, LingBot-Depth, LingBot-Video, LingBot-VLA, and LingBot-VA) to solve specific technical problems, such as aligning modalities, predicting dynamics, and ensuring load balancing in MoE architectures, rather than attempting a single end-to-end model prematurely. Shen introduced the concept of 'embodied-native' models, requiring data native to physical world tasks, model capabilities suited for real-time interaction, and the ability to train from scratch, citing difficulties in replicating self-supervised vision models like DINO and implementing causal modeling for one-way attention. On safety, he stressed that 'fence-based' safety (restricting actions) is insufficient, advocating for 'native safety' embedded during pretraining, similar to human instincts. He also predicted that the 'ChatGPT moment' for embodied AI will occur when ordinary people can easily participate in data collection, making robots truly beneficial and integrated into daily life.
- Abbreviations
- WAIC = World Artificial Intelligence Conference — Всемирная конференция по искусственному интеллекту
- VLA = Vision-Language-Action — Видение-Язык-Действие
- VA = Vision-Action — Видение-Действие
- MoE = Mixture of Experts — Смесь экспертов
- CLIP = Contrastive Language-Image Pre-training — Контрастивное предобучение язык-изображение
- DINO = DIstillation with NO labels — Дистилляция без меток
- GPT = Generative Pre-trained Transformer — Генеративный предобученный трансформер
Source: InfoQ 中国 —
original
