Timeline

Alibaba releases Qwen2.5-Max

Unlike most of Alibaba's Qwen line, Max was released as a proprietary API-only model, pretrained on over 20 trillion tokens, which Alibaba said beat DeepSeek-V3 on several benchmarks.

  • Models & capabilities
  • Benchmarks & progress
  • Minor

Alibaba released Qwen2.5-Max, a large mixture-of-experts model pretrained on more than 20 trillion tokens and refined with supervised fine-tuning and reinforcement learning from human feedback. It arrived nine days after DeepSeek released R1 and amid a broader run of releases from Chinese labs that January, with Alibaba explicitly benchmarking the new model against DeepSeek’s.

Alibaba reported Qwen2.5-Max outperforming DeepSeek-V3 on Arena-Hard, LiveBench, LiveCodeBench and GPQA-Diamond, and scoring competitively against GPT-4o and Claude 3.5 Sonnet on MMLU-Pro — positioning it, by the company’s own comparisons, among the strongest models available at the time on general capability benchmarks, though these were self-reported figures rather than independent evaluation.

Unlike most models in Alibaba’s Qwen family, which the company had released with open weights, Qwen2.5-Max was API-only, accessible through Alibaba Cloud Model Studio or the Qwen Chat interface but not available to download or run locally. The release illustrated a split that widened through 2025: Chinese labs continued to lead much of the open-weight ecosystem while reserving their largest, most capable models for proprietary access — the same trade-off Western labs had made with their own flagship models.