Timeline

Llama 4 lands badly

Meta's mixture-of-experts release was undercut by accusations that a version tuned for LMArena differed from the public weights.

  • Open weights & ecosystem
  • Benchmarks & progress
  • Models & capabilities
  • Major

Meta released two members of its Llama 4 family, Scout and Maverick, as open-weight downloads, alongside benchmark claims for a third, larger model, Behemoth, still in training. All three used a mixture-of-experts design, in which only a fraction of the model’s total parameters activate for any given input: Scout combined 17 billion active parameters across 16 experts (109 billion total), Maverick used the same 17 billion active parameters spread across 128 experts (400 billion total), and Behemoth was described as reaching roughly 288 billion active parameters. Meta said Maverick beat GPT-4o and Gemini 2.0 Flash on several benchmarks, and that Behemoth outperformed GPT-4.5, Claude Sonnet 3.7 and Gemini 2.0 Pro on a number of STEM evaluations.

Within days, the release was overtaken by a dispute about how those claims had been generated. Meta had submitted a version called “Llama-4-Maverick-03-26-Experimental” to LMArena, the crowd-sourced leaderboard that ranks models by human preference in head-to-head comparisons, where it placed second behind only an experimental Gemini 2.5 Pro build. That submitted version produced longer, more stylistically embellished, emoji-heavy responses than the model Meta actually released for download, which users found markedly more concise and, on the same leaderboard, ranked far lower once retested. LMArena said publicly that “Meta’s interpretation of our policy did not match what we expect from model providers,” and updated its submission rules to require that leaderboard entries match publicly released weights. Separately, some users reported the model performing inconsistently across different hosting platforms, which Meta attributed to inference-configuration differences rather than the model itself; other allegations that Llama 4 had been trained on benchmark test data were denied by the company.

The episode became a reference point in arguments about how much leaderboard rankings could be trusted at all, at a moment when LMArena scores were widely used across the industry as a proxy for model quality. Meta’s then chief AI scientist, Yann LeCun, later acknowledged after leaving the company that the results had been “fudged a little bit.”