Alibaba releases Qwen2.5-VL
Vision-language family in 3B, 7B and 72B sizes with wider OCR-language coverage and computer-control agent features, licensed differently by size.
- Open weights & ecosystem
- Models & capabilities
- Colour
Alibaba’s Qwen team released Qwen2.5-VL, an open-weight vision-language model family in 3B, 7B and 72B sizes (a 32B variant followed in March). It continued the Qwen2.5 line with a multimodal encoder aimed at document parsing, video and early computer-use agents.
Alibaba claimed the model extended optical character recognition to 32 languages, up from 10 in the prior generation, with better handling of low light, blur and tilted scans — aimed squarely at real-world document photographs rather than clean scans. It also claimed improved temporal grounding for video understanding and the ability to operate mobile and desktop interfaces as an agent, clicking and typing based on what it sees on screen.
Licensing split by size: the 7B model shipped under Apache 2.0, while the 3B and 72B variants used Alibaba’s own, more restrictive Qwen licence. The release added to a fast-growing field of open-weight vision-language models from Chinese labs competing on cost and permissiveness rather than raw benchmark leadership, part of a broader pattern in which open Chinese releases pressured Western labs’ pricing and licensing choices through 2025.