Google DeepMind says Gemini Robotics 2 is designed to give robots whole-body control, fine manipulation abilities and the capacity to collaborate on complex tasks. The company positions the system as an intelligence layer for adaptable robots operating in less predictable environments.
Google DeepMind has introduced Gemini Robotics 2, an AI system the company says is intended to help robots coordinate whole-body movement, perform fine-grained manipulation and work alongside other robots on complex tasks.
In its announcement, Google DeepMind describes the model as an “intelligence layer” for robots of different shapes and sizes. The company argues that many existing robots are programmed or remotely operated for narrow, repetitive workflows, and that transferring learned capabilities between different robot bodies remains difficult.
Gemini Robotics 2 is presented as an effort to make robot behavior more adaptable in environments that are not tightly controlled. According to Google DeepMind, the system can reason through movements across a robot’s body rather than focusing only on a single arm or gripper.
Google DeepMind says the model is intended to support actions including walking, crouching, stretching and manipulating objects. The company gives the example of a humanoid robot cleaning a cluttered room, a task that would require navigation, changing posture and object handling in sequence.
The emphasis on whole-body control matters because many real-world jobs require robots to balance movement with manipulation. A machine may need to reach into a confined space, reposition itself, handle objects with care and continue monitoring its surroundings. Google DeepMind says Gemini Robotics 2 is designed to bring those capabilities together.
The company also highlights advanced dexterity, suggesting that the model is meant to help robots manage more precise physical interactions than conventional task-specific systems.
Related Google materials describe Gemini Robotics ER 2 as a vision-language model for spatial, temporal and physical reasoning. Google says ER 2 is designed for high-level robot functions including multi-step planning, tracking progress through video and collaboration among multiple robots.
Google DeepMind’s ER 2 model card also documents limitations and distribution details for the vision-language model. That documentation is relevant because robot systems must interpret visual scenes, understand task progress and make decisions around physical action—areas where errors can have consequences beyond those of a conventional chatbot.
Google DeepMind has not, in the provided materials, detailed broad commercial availability, independent performance benchmarks or deployments of Gemini Robotics 2. The announcement instead frames the release as a step toward robots that can learn and adapt more effectively across bodies and changing environments.
For robotics developers, the central claim is not simply that a robot can execute a scripted movement. It is that an AI model can help coordinate perception, planning, locomotion and manipulation as parts of one physical task.
In its announcement, Google DeepMind describes the model as an “intelligence layer” for robots of different shapes and sizes.
The company argues that many existing robots are programmed or remotely operated for narrow, repetitive workflows, and that transferring learned capabilities between different robot bodies remains difficult.
Gemini Robotics 2 is presented as an effort to make robot behavior more adaptable in environments that are not tightly controlled.
Continue reading