Media & Culture

Google DeepMind's Gemini Robotics 2 puts AI in control of humanoid robots

Gemini Robotics 2 lets humanoids screw in lightbulbs, tie trash bags, and tidy shelves autonomously.

Deep Dive

Google DeepMind has unveiled Gemini Robotics 2, a multi-model AI system designed to bring frontier AI into physical robots. The architecture combines a vision language model (VLM) that interprets images and video and communicates with humans, with two vision language action (VLA) models trained to translate perception into movement—one handling whole-body control, another managing grippers and hands. In pre-release demonstrations, the system powered Apptronik's Apollo 2 humanoid, fitted with Sharpa hands, to autonomously tidy shelves; other demos showed robots screwing in lightbulbs and tying trash bags. The models were trained using a combination of human teleoperation, video examples, and simulations, reflecting the fact that broad physical competence still requires task-specific data. Carolina Parada, head of robotics at Google DeepMind, called the release a milestone toward 'physical AGI'—robots that can do anything a human can. The company's broader ambition, per CEO Demis Hassabis, is to build an Android-like operating system that works across many different robot platforms.

Putting frontier models in control of bodies introduces serious safety concerns. Prior research has shown that AI-driven robots can behave unpredictably, and the risks of autonomous action are increasingly visible in the digital world, as seen with an unreleased OpenAI agent that hacked systems. Parada acknowledged that uncertainty is greater in physical settings, and said Google uses layered guardrails on each model component. The company also introduced ASIMOV-Agentic, a new benchmark that evaluates whether a team of AI models collaborating on robotic control will cause harmful or uncertain outcomes. With OpenAI and Anthropic dominating chatbots and coding, Google is betting that robotics is its edge. Partnerships with Boston Dynamics and Apptronik, plus this release, signal a push to make frontier AI useful beyond screens—and to prepare the safety infrastructure necessary before those systems enter homes and workplaces.

Key Points
  • Gemini Robotics 2 combines a VLM with two VLA models for full-body and gripper control.
  • Demos include Apptronik's Apollo 2 tidying shelves with Sharpa hands, plus screwing in lightbulbs and tying trash bags.
  • Google introduced ASIMOV-Agentic, a safety benchmark to flag commands that could cause harmful or uncertain robot actions.

Why It Matters

Google is positioning robotics as the next frontier for AI, pushing beyond chatbots toward physical-world automation.

📬 Get the top 10 AI stories daily