Timeline

Hugging Face retires the Open LLM Leaderboard

The leaderboard had ranked more than 13,000 open models over roughly two years; Hugging Face said fixed multiple-choice tests no longer distinguished reasoning models.

  • Benchmarks & progress
  • Open weights & ecosystem
  • Minor

Hugging Face retired the Open LLM Leaderboard, the fixed-benchmark ranking that had become a standard reference point for the open-weight model ecosystem since it launched, having evaluated more than 13,000 models across two iterations over roughly two years. The team said the leaderboard’s static suite of multiple-choice and knowledge tests could no longer meaningfully distinguish between models as capabilities shifted toward reasoning and assistant-style interaction — the kind of behaviour that mattered to users was no longer well captured by the questions the leaderboard asked.

Hugging Face argued that keeping the leaderboard running risked encouraging developers to keep optimising for benchmarks that had become disconnected from real capability differences, a version of Goodhart’s law playing out over the specific case of open-model evaluation. Rather than replace it with a single successor, the company pointed to a fragmented ecosystem of more than 200 community-run, domain-specific leaderboards on its platform, covering areas such as mathematics and coding, and said its own team would continue evaluation work under a separate “Open Evals” effort.

The retirement was a visible acknowledgment that a widely cited, single-number ranking of open models had run its course, a milestone in the broader move away from static academic benchmarks toward more targeted or dynamic evaluation as reasoning models made older test suites easier to saturate or game.