Timeline

Anthropic tests Claude on BioMysteryBench

On 23 questions its own expert panel could not solve, an unreleased preview model Anthropic called Mythos scored roughly 30%, against single digits for Claude Haiku 4.5.

  • Benchmarks & progress
  • Minor

Anthropic published BioMysteryBench, a set of 99 bioinformatics research questions written by domain experts, designed to test whether Claude models could carry out the kind of open-ended, method-agnostic analysis real bioinformatics research requires, rather than the closed-form knowledge and reasoning questions typical of existing science benchmarks. Tasks drew on real datasets — DNA and RNA sequencing, proteomics, metabolomics — with access to a minimal set of canonical tools, public databases such as NCBI and Ensembl, and the ability to install further software.

Anthropic split the questions by whether at least one member of its expert panel could solve them: 76 were human-solvable, establishing a difficulty baseline, and 23 were not, which Anthropic said might reflect problems genuinely beyond current human capability or simply extreme difficulty rather than impossibility — a distinction it said the benchmark could not resolve on its own. An unreleased preview model Anthropic referred to as Mythos scored 82.6% on the solvable set and 29.6% on the human-difficult set, compared with 36.8% and 5.2% for the smaller, publicly available Claude Haiku 4.5.

Anthropic cautioned that performance on the hardest questions was inconsistent: repeating the same difficult task five times, the model often succeeded only once or twice, suggesting it was sometimes stumbling onto a workable approach rather than following a repeatable method. Transcript analysis pointed to two things separating its approach from a human researcher’s — drawing on a much larger span of published literature to synthesise an answer without running new analysis, and layering multiple independent methods together when uncertain. Anthropic framed the benchmark as a way to measure whether a model could contribute to open research problems, rather than only recall or reason about already-established science.