Google DeepMind Unveils Gemini Robotics 2 to Advance 'Physical AGI'
A new three-model intelligence layer enables humanoid robots to reason through complex tasks and adapt to new hardware with minimal data.
Google DeepMind has released Gemini Robotics 2, a sophisticated intelligence layer designed to move the industry toward "physical AGI." The system aims to create generalist robots capable of performing any task a human can, bridging the gap between high-level cognitive reasoning and low-level physical execution.
The system is built on a three-model architecture consisting of a primary Vision-Language-Action (VLA) model, an Embodied Reasoning engine (ER 2), and a low-latency On-Device 2 model. The ER 2 model can identify key moments in video feeds with nearly 90% accuracy and classify video frame completeness with almost 60% accuracy. To ensure safety, Google introduced the ASIMOV-Agentic benchmark, which evaluates whether agents correctly refuse unsafe tool calls or request human assistance. The system has already undergone testing on Boston Dynamics hardware.
The Shift Toward Generalist Robotics
For years, the robotics field has been constrained by narrow, pre-programmed scripts or the need for constant human teleoperation. By integrating Vision-Language Models (VLMs) with VLAs, DeepMind is attempting to move away from these rigid constraints. This approach allows robots to perceive their environment and reason through movements rather than simply following a fixed set of instructions. This transition represents a fundamental change in how machines interact with the physical world, moving from execution to understanding.
Implications for Hardware Deployment
This architecture represents a significant shift toward robots that can perceive, reason, and correct their own mistakes in real-time. A critical advantage is the flexibility of the On-Device 2 model, which can adapt to entirely new robot designs using only about 200 examples or a few hours of movement data. This efficiency drastically lowers the barrier to deploying AI across diverse platforms, as the intelligence layer remains consistent even when the physical chassis changes.
Next Steps for Developers
While the full system is being refined, the Embodied Reasoning (ER 2) model is currently available to developers through the Gemini Live API. Future developments will likely focus on further improving the accuracy of video frame classification and expanding the range of hardware the system can control without extensive retraining. As these models evolve, the goal remains a seamless integration of digital intelligence and physical dexterity, potentially redefining industrial automation and personal assistance.