ModelsRobotics 🇺🇸 30.07.2026 18:01

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Google/DeepMindGoogle/DeepMind Boston DynamicsBoston Dynamics
Google DeepMind has launched Gemini Robotics ER 2, a new embodied reasoning model for robotics. It enables real-time video understanding, multi-step task planning, and multi-robot collaboration, and is available via the Gemini API and Google AI Studio.
Google DeepMind introduced Gemini Robotics ER 2, an embodied reasoning model for robotics that serves as a high-level brain for robots. It can chat with humans, understand the physical world, plan multi-step tasks, and hand off motor execution to lower-level vision-language-action (VLA) models. The model natively calls tools like Google Search or user-defined functions. It improves over ER 1.6 by processing continuous video feeds to track progress, adapt to failures, and determine task completion. Multi-robot collaboration is introduced, enabling diverse robots to work together. Gemini Robotics ER 2 achieves 57.4% accuracy on progress classification and 91.3% accuracy on moment-finding tasks with sub-second latency. It outperforms previous models on safety benchmarks, including halting a humanoid robot when a person is nearby. The model is publicly available via the Gemini API, Google AI Studio, and in private preview on Gemini Enterprise Agent Platform.
Сокращения
VLA = Vision-Language-Action — Vision-Language-Action
API = Application Programming Interface — интерфейс программирования приложений
Source: Google DeepMind — original
Our earlier posts on this topic ↓
Fresh news