Timeline

DeepMind's SIMA 2 uses Gemini to reason and act inside 3D game worlds

The agent, released as a limited research preview, generalised to games and AI-generated worlds it had not been trained on.

  • Models & capabilities
  • Minor

Google DeepMind released SIMA 2, a successor to its Scalable Instructable Multiworld Agent, as a limited research preview open to a small cohort of academics and game developers. Where the original SIMA followed short, literal instructions inside 3D game environments, SIMA 2 uses Gemini as its reasoning core, trained on human demonstration videos combined with Gemini-generated language labels, so the agent can explain its intentions in language rather than acting as a black box.

DeepMind reported that SIMA 2 completed a substantially higher share of tasks than SIMA 1 across the environments it was trained on, and that it also performed well in games and AI-generated worlds it had never encountered during training, including MineDojo and ASKA. The agent handles multimodal input — text, sketches, different languages and emojis — and DeepMind described interacting with it as closer to collaborating with a companion than issuing commands, since it can improve at a task through trial and error using Gemini’s feedback.

DeepMind framed SIMA 2 as a step toward general-purpose embodied intelligence, arguing that skills such as navigation and tool use learned in virtual worlds should transfer to physical robotics. That claim was not independently tested at release, and the restricted preview meant outside researchers could not yet verify DeepMind’s completion-rate figures against the original SIMA. The release nonetheless extended a pattern among frontier labs of using game environments as a proving ground for agentic reasoning ahead of deployment in messier, real-world settings.