LMArena responds to 'Leaderboard Illusion' paper with policy changes
LMArena disputed the paper's headline figures on open-model share and score-boosting but agreed to mark scores 'provisional' and disclose pre-release testing.
- Benchmarks & progress
- Minor
LMArena, which runs the crowdsourced Chatbot Arena leaderboard, published a formal response to “The Leaderboard Illusion,” an academic critique co-authored by researchers at Cohere Labs, AI2, Princeton and Stanford that argued undisclosed private pre-release testing let large labs game their rankings.
LMArena disputed several of the paper’s central figures. Against the paper’s claim that open-weight models held only 8.8% of leaderboard share, LMArena cited its own count of 40.9%, arguing the authors had missed models such as Llama and Gemma. It also rejected the paper’s estimate that private testing could inflate scores by more than 100 Elo points, calling the underlying simulation “arbitrarily chosen” and unrelated to real Arena data; its own figures put a typical boost at around 11 points after 50 test rounds, fading as public voting data accumulated. It said its policy of permitting multiple pre-release model variants had been public since March 2024.
Despite the disputes, LMArena said it would make several changes: publishing which submitted models are pre-release variants and when they are retired from testing, and labelling a model’s score “provisional” until it has accumulated 2,000 fresh votes if more than ten variants were tested privately in parallel before release.
The exchange followed an earlier controversy in which Meta was accused of tuning a special Llama 4 Maverick variant to rank higher than the model it actually shipped, and it added to a broader argument about whether crowdsourced preference leaderboards measured what they claimed to.