Timeline

ASI-Bench measures how far agents are from autonomous science

Withdrawing step-by-step guidance and forcing agents to choose their own research methods dropped the average score from 50.91 to 26.62 on a 0–100 scale.

  • Benchmarks & progress
  • Minor

A 42-author team spanning institutions including Tsinghua University, MIT, Harvard, Carnegie Mellon and Microsoft Research published ASI-Bench, a benchmark built to test whether AI agents can carry out complete, project-level scientific research rather than answer narrow, self-contained questions. The benchmark comprises 60 tasks across 11 fields — including mathematics, physics, chemistry, biology, medicine and computer science — built by more than 40 domain experts at a cost the authors put at over 31,000 human hours, then checked through expert review, AI-assisted auditing, sandboxed execution and scorer validation.

ASI-Bench’s distinguishing feature is that it progressively withdraws human methodological guidance within the same research project, rather than testing a fixed task once. Across 18 agent-model configurations, the authors reported an average score of 50.91 out of 100 when agents were given a detailed procedure to follow, falling to 29.10 when the procedure was removed but the method to use remained specified, and to 26.62 when agents had to determine the research method themselves as well as carry it out.

The authors concluded that current systems “remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research” — a more cautious reading than the benchmark’s title might suggest, and one that lands alongside other August 2026 results, including London startup Inherent’s Faraday model, pointing at how far current agents remain from independent scientific work even as narrower AI-for-science tools mature.