Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Google/DeepMind
Boston Dynamics
Google DeepMind has launched Gemini Robotics ER 2, a new embodied reasoning model for robotics. It enables real-time video understanding, multi-step task planning, and multi-robot collaboration, and is available via the Gemini API and Google AI Studio.
Google DeepMind introduced Gemini Robotics ER 2, an embodied reasoning model for robotics that serves as a high-level brain for robots. It can chat with humans, understand the physical world, plan multi-step tasks, and hand off motor execution to lower-level vision-language-action (VLA) models. The model natively calls tools like Google Search or user-defined functions. It improves over ER 1.6 by processing continuous video feeds to track progress, adapt to failures, and determine task completion. Multi-robot collaboration is introduced, enabling diverse robots to work together. Gemini Robotics ER 2 achieves 57.4% accuracy on progress classification and 91.3% accuracy on moment-finding tasks with sub-second latency. It outperforms previous models on safety benchmarks, including halting a humanoid robot when a person is nearby. The model is publicly available via the Gemini API, Google AI Studio, and in private preview on Gemini Enterprise Agent Platform.
- Сокращения
- VLA = Vision-Language-Action — Vision-Language-Action
- API = Application Programming Interface — интерфейс программирования приложений
Source: Google DeepMind —
original
