Timeline

Google DeepMind's Co-Scientist AI is validated in physical lab experiments

The system designed a chemical-vapour-deposition route that produced a two-dimensional material on a physical reactor, beyond the hypothesis-only role of earlier versions.

  • Models & capabilities
  • Benchmarks & progress
  • Notable

Samuel Schmidgall led a 35-author team, credited to Google DeepMind and Google Research, that posted a paper describing an extended version of Co-Scientist — the Gemini-based multi-agent research system DeepMind had previously shown generating hypotheses for human scientists to test — built on the Gemini 3 Deep Think model. Unlike earlier versions, this account reported results reaching execution on physical laboratory hardware rather than stopping at a proposal for humans to carry out.

In materials science, the system designed a precursor route for a class of two-dimensional materials called MXenes and had it run on a chemical vapour deposition reactor, producing what the authors described as a lamellar 2D material sharing key structural similarities with a target MXene lattice — a result they caveated as still requiring confirmation of its exact atomic structure. The paper also reported the system predicting swarming behaviour in engineered E. coli bacteria from sparse imaging data, matching what the authors said were unpublished wet-lab measurements; an inference-time scaling method the system itself discovered, which the authors said outperformed six frontier models on the hard and professional subsets of the HealthBench benchmark while reducing potential clinical harm in a physician evaluation; and a double-blind study in which 30 domain experts conducted 450 reviews of AI-generated papers, testing whether built-in reliability checks reduced hallucination and plagiarism.

The paper marks a step beyond DeepMind’s May 2026 Nature publication on Co-Scientist, whose six case studies had human researchers execute AI-proposed experiments and confirm the results themselves; here, the system’s own outputs drove the physical apparatus directly. As with the earlier work, the case studies are DeepMind’s own selected examples rather than an independent trial, and the paper had not been through peer review at the time it was posted.

In the commentary

What people were saying around this time — external links, from the record's commentary rail.