Timeline

Baidu launches ERNIE 5.0, a 2.4-trillion-parameter native multimodal model

Baidu said the mixture-of-experts model activates under 3% of its parameters per query and ranked first among Chinese models, eighth globally, on LMArena's text leaderboard.

  • Models & capabilities
  • Benchmarks & progress
  • Notable

Baidu launched the full version of ERNIE 5.0, a 2.4-trillion-parameter mixture-of-experts model, having previewed it two months earlier. The mixture-of-experts architecture activates fewer than 3% of its parameters on any given query, keeping inference costs down despite the model’s overall size. Baidu described it as “natively multimodal,” trained from the outset to handle text, images, audio and video together rather than bolting separate modality encoders onto a text-only base.

On LMArena’s leaderboards, Baidu said ERNIE 5.0 ranked first among Chinese-developed models and eighth globally on text tasks, ahead of OpenAI’s GPT-5.1-High on that measure — though LMArena rankings reflect crowdsourced human preference votes on chat responses rather than a fixed benchmark suite, and are one of several competing ways labs claim leadership. The launch came alongside Baidu’s disclosure that its Ernie chatbot and assistant products had reached 200 million monthly active users.

The release continued a pattern in which Chinese labs — Baidu, Alibaba, DeepSeek, Moonshot and others — released large frontier-scale models on a roughly quarterly cadence through 2025 and into 2026, closing much of the gap in benchmark performance with US labs while operating under continued US export restrictions on the most advanced AI training chips. ERNIE 5.0’s scale and multimodal design put it in direct competition with both domestic rivals and the largest models from OpenAI, Google and Anthropic on general capability, even as questions remained about how closely LMArena preference rankings tracked capability on harder, more verifiable tasks.