Timeline

Tencent open-sources Hunyuan Image 3.0

An 80-billion-parameter mixture-of-experts model, trained on 5 billion image-text pairs, released under a licence that excludes the EU, UK and South Korea.

  • Open weights & ecosystem
  • Models & capabilities
  • Minor

Tencent released the weights of Hunyuan Image 3.0, an image-generation model it described as the first open-source model of commercial grade built as a “native multimodal” system rather than a separate image generator bolted onto a text model. The model used a mixture-of-experts architecture combining 64 experts with a Transfusion-style unified framework, totalling around 80 billion parameters with roughly 13 billion active on any given generation — a departure from the diffusion-transformer designs common among other open image models. Tencent reported it was trained on around 5 billion image-text pairs and 6 trillion text tokens, and said the model was capable of rendering long passages of text accurately within images and drawing on broader world knowledge to interpret complex prompts.

Running the model required substantial hardware — Tencent’s own guidance called for at least three 80GB GPUs, with four recommended, alongside 170GB of storage and 64GB or more of system memory — putting it well beyond consumer hardware and closer to the resources needed for the large open text models Chinese labs had already been releasing.

The weights were published under Tencent’s own Hunyuan Community Licence, which permitted commercial use and modification but excluded the European Union, United Kingdom and South Korea from that permission, and required products with more than 100 million monthly active users to obtain a separate licence from Tencent. The release added to the run of open-weight image and video models coming out of Chinese labs through 2025, extending the open-versus-closed contest that had mostly played out in text models into image generation.