Timeline

Zhipu reveals GLM-5.3-Flash after testing it anonymously online for a week

Priced at roughly a tenth of GLM-5.2 and released under an MIT licence, it is the first natively multimodal model in the GLM-5 series.

  • Models & capabilities
  • Open weights & ecosystem
  • Notable

Zhipu AI, trading as Z.ai, released GLM-5.3-Flash, a 320-billion-parameter model with 18 billion active per token, under an MIT licence and priced, the company said, at roughly a tenth of its predecessor GLM-5.2. Its model card describes it as the first natively multimodal model in the GLM-5 series and the first in the line to combine sparse and linear attention, an architecture change aimed at cutting the cost of serving long-context requests.

The release also confirmed the identity of “Ox Alpha,” a model that had run anonymously on OpenRouter and the coding tool OpenCode from 20 August without Z.ai disclosing itself as the developer. Outside researchers had narrowed down the likely source before the official reveal by comparing its tokenizer behaviour with known GLM models; Z.ai’s own announcement then confirmed it had tested GLM-5.3-Flash under that stealth listing to gauge real-world demand ahead of a named release — a tactic several model developers have used on public routing platforms such as OpenRouter.

Z.ai’s model card said GLM-5.3-Flash was “approaching” Claude Opus 4.8 on coding and agentic benchmarks, a comparison the company made itself rather than one independently verified. The release came twelve days after Zhipu shipped the full-size GLM-5.3 as an API-only product with open weights promised within about two weeks; those larger weights had still not been published when Flash shipped, leaving the smaller model as the only open-weight release of the GLM-5.3 generation.