Alibaba releases Qwen2.5 model family
Alibaba's release spanned seven sizes from 0.5B to 72B parameters, plus dedicated coding and maths variants, trained on 18 trillion tokens.
- Open weights & ecosystem
- Models & capabilities
- Notable
Alibaba’s Qwen team released Qwen2.5, a family of large language models spanning seven sizes from 0.5B to 72B parameters, alongside dedicated Qwen2.5-Coder and Qwen2.5-Math variants. The base models were pretrained on what Alibaba described as 18 trillion tokens, and the coder models on a further 5.5 trillion tokens of code-specific data.
All sizes except the 3B and 72B were released under the Apache 2.0 licence, making most of the family freely usable, including commercially, while the largest and one small model remained under a more restrictive Qwen licence. Alibaba reported the 72B instruction-tuned model scoring above 85 on MMLU, above 85 on HumanEval and above 80 on MATH, and said it competed with proprietary systems on these benchmarks; it also said even the much smaller 3B model performed competitively for its size.
The release arrived amid a wave of open-weight model families — Meta’s Llama, Mistral’s line-up and Alibaba’s own earlier Qwen2 — competing on the argument that openly available weights could match closed, proprietary models like GPT-4o on standard benchmarks while giving developers the ability to fine-tune, self-host and inspect what they were running. Qwen2.5’s breadth, from a 0.5B model runnable on modest hardware to a 72B flagship, made it one of the most complete open releases to date, and its coding and maths specialists in particular were widely adopted by developers building on open models through the rest of 2024.