Why It's Hard for a Humanoid to Grasp Moving Objects: Developing VLA for Dynamic Environments
Unitree Robotics
ML engineer Daniya from MTC Web Services discusses the challenges VLA models face when dealing with moving objects, and how the company's R&D unit adapts the RLDX-1 architecture for the Unitree G1 humanoid to pick parts from a conveyor belt while maintaining balance.
An ML engineer from MTC Web Services, Daniil, explains why humanoid robots struggle to operate in dynamic environments, such as when objects move along a conveyor belt. In static conditions, modern VLA (Vision-Language-Action) models demonstrate high control quality, but in dynamic settings, three main issues arise: latency between observation and action (camera image capture, processing, and command transmission take tens of milliseconds, during which the object shifts), obsolescence of precomputed action chunks, and insufficient information from analyzing a single frame (it is impossible to determine the speed and direction of motion). To address these problems, models are transitioning to analyzing frame sequences, continuously updating actions, and incorporating physical signals (joint positions, tactile data). A humanoid is more complex than a stationary manipulator, because arm movement affects body balance; whole-body control is required, coordinating the movements of arms, torso, and legs. At RnD MWS, they are developing a whole-body dynamic VLA based on the RLDX-1 architecture, which processes multiple streams of information (motion, context, physics, actions) and works with the Unitree G1 humanoid robot. The goal is to teach the robot to grasp parts on a conveyor belt by predicting their trajectory, synchronizing movements, and maintaining balance.
Source: Habr — хаб ИИ —
original
