Google DeepMind has made a significant leap in the field of robotics with the release of Gemini Robotics 2, a suite of three advanced AI models designed to enhance robot intelligence and physical capabilities. This new system represents a major step forward in the development of autonomous robots that can navigate complex environments and perform sophisticated tasks.
Three Key Models for Advanced Robotics
The first model, a vision-language-action system, enables whole-body control for humanoid robots. This model allows robots to process visual inputs, understand language, and execute actions seamlessly across their entire body. The second model, Gemini Robotics ER 2, focuses on embodied reasoning and task orchestration, enabling robots to understand and plan complex tasks in real-world settings. The third model, an on-device Vision-Language-Action (VLA) system, is capable of adapting to new robot bodies within hours, significantly reducing the time and effort required for robot customization.
Real-World Deployment and Public Availability
These models are already being deployed in physical robots such as the Apptronik Apollo 2 and the Franka Duo. While the first two models are being used internally, only the ER 2 model is publicly available, offering researchers and developers a glimpse into the future of robot intelligence. This release underscores DeepMind's commitment to advancing robotics through AI-driven solutions.
Implications for the Future of Robotics
The introduction of these models signals a shift toward more intelligent, adaptable, and collaborative robots. As AI continues to evolve, such advancements will likely pave the way for robots that can operate autonomously in dynamic environments, work alongside humans, and perform increasingly complex tasks. With these new tools, DeepMind is positioning itself at the forefront of the robotics revolution, pushing the boundaries of what is possible in physical AI systems.



