'The Leaderboard Illusion' paper critiques Chatbot Arena methodology
Researchers found Meta tested roughly 27 private Llama variants before its public release and that OpenAI and Google alone received about 40% of all Arena battle data.
- Benchmarks & progress
- Notable
A group of thirteen researchers led by Cohere Labs’ Sara Hooker, with co-authors from the Allen Institute for AI, Princeton, Stanford and elsewhere, published a lengthy critique of Chatbot Arena’s ranking methodology, arguing that undisclosed private pre-release testing let a handful of well-resourced labs quietly optimise for the leaderboard before ever appearing on it publicly. The paper reported that Meta had tested roughly 27 private variants of its models ahead of the Llama 4 Maverick release, only later disclosing which configuration would ship, and calculated that limited extra access to Arena battle data could translate into score gains of up to 112% relative to a model’s baseline Arena performance.
The paper also quantified a data-access asymmetry: it found Google and OpenAI’s models alone accounted for roughly 40% of all battles recorded on the platform, while 83 open-weight models combined received under 30%, and argued that proprietary models were also more often retained rather than silently deprecated after weak showings, compounding the advantage.
LMArena disputed several of the paper’s specific claims and characterisations but subsequently announced policy changes around private testing disclosure and score reporting. The episode, arriving weeks after Meta was separately accused of submitting a chat-tuned Llama variant to the Arena that differed from the model it actually released, sharpened a running argument that crowdsourced human-preference leaderboards had become a target labs could game rather than a neutral yardstick, and it fed a broader shift in the field toward benchmarks with harder-to-manipulate methodologies.