Google DeepMind has announced Gemini Robotics 2, a vision-language-action model designed for robot control across systems ranging from tabletop machines to humanoids. The company also introduced Gemini Robotics ER 2, a related model for video understanding, multi-step task coordination and collaboration between robots.
Google DeepMind has introduced Gemini Robotics 2, a vision-language-action, or VLA, model intended to control robots that require more than single-arm manipulation.
In its announcement, Google DeepMind described the release as an effort to bring “whole body intelligence” to robots. The model is paired with Gemini Robotics ER 2, a vision-language reasoning model focused on interpreting visual information and organizing longer tasks.
The releases outline a split between physical robot control and higher-level task reasoning. Gemini Robotics 2 is intended to generate actions for a machine, while ER 2 is positioned to help systems understand scenes, plan work and monitor progress.
Google DeepMind’s Gemini Robotics 2 product page says the VLA model can support robots ranging from tabletop systems to full humanoids. The company describes its target capability as “feet-to-fingertips control,” indicating that it is designed for machines whose tasks involve coordinated movement across an entire body rather than only a robotic gripper.
That scope is relevant to mobile and humanoid machines, which may need to combine navigation, balance, visual perception and object handling to complete a task. A robot asked to retrieve or move an item, for example, may need to travel through an environment, position itself safely and manipulate the object without losing stability.
VLA models combine visual inputs, language instructions and action outputs. In principle, this allows a robot to use what it sees and a user’s request as inputs for physical behavior. Google DeepMind’s materials frame Gemini Robotics 2 as an extension of this approach to a wider range of hardware forms.
The cited product and announcement pages describe intended capabilities, but do not provide benchmark figures for comparing Gemini Robotics 2 with other robotics models.
Google introduced Gemini Robotics ER 2 alongside the control model. According to a Google blog post, ER 2 provides video understanding, multi-step task orchestration, self-correction and collaboration among different robots.
Video understanding could allow a robotic system to interpret changes in a work area or assess whether a step has been completed. Multi-step orchestration is intended to organize a broader task into successive actions, while self-correction suggests a role in checking progress and adjusting when an attempted action does not produce the expected result.
Google also says ER 2 can support collaboration among different robots. Such coordination could be useful where separate machines need to share work, although the company’s cited materials do not detail specific deployments or performance results.
Google DeepMind says Gemini Robotics ER 2 is available through Google AI Studio and in private preview for the Gemini Enterprise Agent Platform. The cited sources do not state that Gemini Robotics 2 has a general public release.
Together, the releases show Google DeepMind pursuing a two-part architecture for robotics: a model for whole-body physical action and another for visual interpretation and task-level organization. The company is positioning the combination for robots that must both act in the physical world and manage more complex instructions over time.
How well the approach transfers across robot hardware and real-world environments remains to be established. For now, Google DeepMind’s published descriptions emphasize the planned range of capabilities: full-body control, video-based understanding, multi-step work, self-correction and multi-robot collaboration.
Google DeepMind expands its robotics models Google DeepMind has introduced Gemini Robotics 2 , a vision language action, or VLA, model intended to control robots that require more than single arm manipulation.
In its announcement, Google DeepMind described the release as an effort to bring “whole body intelligence” to robots.
The model is paired with Gemini Robotics ER 2 , a vision language reasoning model focused on interpreting visual information and organizing longer tasks.
Continue reading