RoboticsModels 🇨🇳 05.08.2026 11:01

Google launches three new models to turn robots into 'all-around workers': quick adaptation, full-body control, and team collaboration

Google/DeepMindGoogle/DeepMind Google DeepMindGoogle DeepMind
Google DeepMind announced Gemini Robotics 2 on July 30, a suite of three AI models that enable whole-body control, advanced dexterity, and multi-robot collaboration for humanoid and other robots. The models include a VLA model for full-body control, an embodied reasoning VLM for planning and collaboration, and a device-local model for rapid adaptation. Google emphasizes safety with a new benchmark called ASIMOV-Agentic.
On July 30, Google DeepMind introduced Gemini Robotics 2, a new version of its AI model series designed to control various robots, including humanoids, enabling them to autonomously adapt to unpredictable environments and reason about each action, thus unlocking a wide range of tasks such as cleaning up trash and picking up a watering can. The update adds intelligent whole-body control, advanced dexterity, and multi-robot collaboration, marking a new phase in robotics. The previous Gemini Robotics system could drive robots to perform fine operations like sealing bags and folding origami, but only with robotic arms and hands. The new models can run on-device and adapt to entirely new robot bodies within hours, transferring learned skills across different platforms. They also allow different robots to work together to complete tasks faster. Google sees this as a milestone toward physical AGI, where robots can do anything a human can do. Demo videos show the Apptronik Apollo 2 humanoid following commands like placing a watering can on a green box, and robots performing tasks such as placing books on shelves, tidying boards, and putting tapes into a radio. The system supports highly dexterous five-finger hands, like the 22-degree-of-freedom SharpaWave hand on Apollo 2, performing tasks like tying knots and sealing zipper bags, as well as standard two-finger grippers like the Franka Duo. Google reports success rates between 45.7% and 76.3% for three whole-body manipulation categories, 74.2% to 89.6% for three Franka Duo gripper categories, and 32% to 92% for multi-finger tasks. The videos show fully autonomous robots, contrasting with Elon Musk's Optimus which reportedly used remote operators. However, these robots are not general-purpose; each task is specifically trained using teleoperation, video examples, and simulation. Google's approach involves three distinct models: Gemini Robotics 2, a VLA model converting visual and language inputs into motor control for full-body control; Gemini Robotics ER 2, an embodied reasoning VLM acting as the brain agent for communication, planning multi-step tasks, and enabling robot team collaboration; and Gemini Robotics On-Device 2, a VLA model for local execution with fast adaptation to new robot bodies, inherited from Gemini Robotics 1.5. For safety, Google introduced ASIMOV-Agentic, a benchmark for agentic safety orchestration and uncertainty resolution, measuring the agent's ability to reject unsafe tool calls and predict task completion. Gemini Robotics ER 2 is the safest robot model yet, detecting human proximity and stopping robots accordingly. The software runs on robot bodies, many of which are made in China, despite recent US restrictions on Chinese-made robots.
Abbreviations
AGI = Artificial General Intelligence — Общий искусственный интеллект
VLA = Vision-Language-Action — Модель «зрение-язык-действие»
VLM = Vision-Language Model — Мультимодальная модель «зрение-язык»
Source: InfoQ 中国 — original
Our earlier posts on this topic ↓
Fresh news