Google DeepMind introduces Gemini Robotics for physical-world tasks
Google DeepMind said the vision-language-action model more than doubled a rival system's score on a generalisation benchmark, folding physical actions into Gemini as a new output type.
- Models & capabilities
- Notable
Google DeepMind introduced Gemini Robotics, a vision-language-action model built on Gemini 2.0 that adds physical robot actions as a new output modality alongside text and images, and a companion model, Gemini Robotics-ER (embodied reasoning), aimed at developers who wanted to pair Gemini’s spatial understanding with their own low-level robot controllers rather than take full end-to-end control.
Google DeepMind described the models as advancing on three fronts: generality across novel objects, environments and instructions the model had not been trained on; interactivity, including following spoken commands in multiple languages and adapting to changes mid-task; and dexterity, demonstrated with tasks such as folding origami and packing a snack into a sealed bag — manipulation skills that had historically challenged learned robot-control systems more than simulated or purely digital tasks. The company reported that Gemini Robotics more than doubled the performance of other state-of-the-art vision-language-action models on a generalisation benchmark, and that Gemini Robotics-ER achieved two to three times the success rate of Gemini 2.0 alone on end-to-end control tasks.
The announcement, released the same day as the open-weight Gemma 3 family, marked Google’s clearest bid to extend a foundation model’s reasoning into physical embodiment rather than leaving robotics to narrower, task-specific systems — a direction rival labs and robotics start-ups pursued through the rest of 2025 with their own vision-language-action releases.
Gemini Robotics-ER was made available to select partners for testing, while Google DeepMind described Gemini Robotics itself as still primarily a research system, tested on the company’s own bimanual robot platforms rather than shipped as a product. As with other vision-language-action announcements in this period, the reported benchmark comparisons were Google’s own; the company did not publish the generalisation benchmark’s full methodology or open it to independent replication at launch, which was typical of robotics announcements at this stage of the field, where standardised, cross-lab evaluation remained less mature than for language benchmarks.