Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Google DeepMind Introduces Gemini Robotics 2 for Whole-Body Robot Control
Kaino
YesterdayJul 30, 2026, 12:00 AM0 views

Google DeepMind Introduces Gemini Robotics 2 for Whole-Body Robot Control

Google DeepMind has announced Gemini Robotics 2, a vision-language-action model designed for robot control across systems ranging from tabletop machines to humanoids. The company also introduced Gemini Robotics ER 2, a related model for video understanding, multi-step task coordination and collaboration between robots.

llmsgeminiDeepMind

Google DeepMind expands its robotics models

Google DeepMind has introduced Gemini Robotics 2, a vision-language-action, or VLA, model intended to control robots that require more than single-arm manipulation.

In its announcement, Google DeepMind described the release as an effort to bring “whole body intelligence” to robots. The model is paired with Gemini Robotics ER 2, a vision-language reasoning model focused on interpreting visual information and organizing longer tasks.

The releases outline a split between physical robot control and higher-level task reasoning. Gemini Robotics 2 is intended to generate actions for a machine, while ER 2 is positioned to help systems understand scenes, plan work and monitor progress.

From tabletop systems to humanoids

Google DeepMind’s Gemini Robotics 2 product page says the VLA model can support robots ranging from tabletop systems to full humanoids. The company describes its target capability as “feet-to-fingertips control,” indicating that it is designed for machines whose tasks involve coordinated movement across an entire body rather than only a robotic gripper.

That scope is relevant to mobile and humanoid machines, which may need to combine navigation, balance, visual perception and object handling to complete a task. A robot asked to retrieve or move an item, for example, may need to travel through an environment, position itself safely and manipulate the object without losing stability.

VLA models combine visual inputs, language instructions and action outputs. In principle, this allows a robot to use what it sees and a user’s request as inputs for physical behavior. Google DeepMind’s materials frame Gemini Robotics 2 as an extension of this approach to a wider range of hardware forms.

The cited product and announcement pages describe intended capabilities, but do not provide benchmark figures for comparing Gemini Robotics 2 with other robotics models.

ER 2 focuses on visual reasoning and coordination

Google introduced Gemini Robotics ER 2 alongside the control model. According to a Google blog post, ER 2 provides video understanding, multi-step task orchestration, self-correction and collaboration among different robots.

Video understanding could allow a robotic system to interpret changes in a work area or assess whether a step has been completed. Multi-step orchestration is intended to organize a broader task into successive actions, while self-correction suggests a role in checking progress and adjusting when an attempted action does not produce the expected result.

Google also says ER 2 can support collaboration among different robots. Such coordination could be useful where separate machines need to share work, although the company’s cited materials do not detail specific deployments or performance results.

Google DeepMind says Gemini Robotics ER 2 is available through Google AI Studio and in private preview for the Gemini Enterprise Agent Platform. The cited sources do not state that Gemini Robotics 2 has a general public release.

A paired approach to embodied AI

Together, the releases show Google DeepMind pursuing a two-part architecture for robotics: a model for whole-body physical action and another for visual interpretation and task-level organization. The company is positioning the combination for robots that must both act in the physical world and manage more complex instructions over time.

How well the approach transfers across robot hardware and real-world environments remains to be established. For now, Google DeepMind’s published descriptions emphasize the planned range of capabilities: full-body control, video-based understanding, multi-step work, self-correction and multi-robot collaboration.

Key takeaways
  • 1

    Google DeepMind expands its robotics models Google DeepMind has introduced Gemini Robotics 2 , a vision language action, or VLA, model intended to control robots that require more than single arm manipulation.

  • 2

    In its announcement, Google DeepMind described the release as an effort to bring “whole body intelligence” to robots.

  • 3

    The model is paired with Gemini Robotics ER 2 , a vision language reasoning model focused on interpreting visual information and organizing longer tasks.

Continue reading

Latest from Kaino News

Story pulse

Freshness

Yesterday

Views

0

Reading

3 min

Byline

Kainotomic Team

Utilities

Topics

llmsgeminiDeepMind

Sources

Reference material and original reporting used in this story.

Google DeepMind

Published Jul 30, 2026, 12:00 AM

View source
Google DeepMind Introduces Gemini Robotics 2 for Whole-Body Robot Control · News · Kaino