Timeline

Benchmark finds AI agents complete about a fifth of full science workflows

The 97 released tasks span six domains, from quantum chemistry to life science; a high partial score did not reliably mean a task was actually finished.

  • Benchmarks & progress
  • Minor

A 16-author team, whose institutional affiliation the paper does not state, published FrontierChallenge, a benchmark of 300 end-to-end scientific workflows spanning six domains — quantum chemistry, molecular dynamics, materials characterisation, analytical chemistry, life science, and electrochemistry and environmental science — of which the authors released and evaluated 97 in this paper. The authors argued that existing agent benchmarks mostly score a final answer, an isolated program, or a single narrow domain, leaving open how well an agent can carry a realistic scientific project from data through to a finished, checkable result.

Testing twelve frontier models across three different agent scaffolds, the authors reported that the best-performing combination completed only 20 of the 97 released tasks end to end, a pass rate of 20.6%. They found that neither a high partial score on a task nor a model’s own confident claim to have finished it reliably indicated the work was actually complete to a usable standard — models frequently reported success on workflows that, checked against the paper’s own criteria, had not in fact been finished.

The result follows, by a week, two other agent benchmarks that reached similarly cautious conclusions from different angles: ASI-Bench, which found scores collapsed once agents had to choose their own research method rather than follow a specified one, and StartupBench, which found a comparable shortfall on real commercial workflows rather than scientific ones. None of the three treats low completion rates as evidence against agents’ usefulness on narrower sub-tasks, only against present claims of full end-to-end autonomy.