Timeline

Z.ai says GLM-5.3 built much of its own inference stack

The company said an agent built on its own model tripled inference throughput on a 100,000-chip domestic cluster within two weeks, while insisting humans kept control of objectives and risk.

  • Models & capabilities
  • Compute & infrastructure
  • Minor

Z.ai (Zhipu) said an “Infra Agent” built on its own GLM-5.3 model carried out much of the work of standing up a production inference system for GLM-5.3-Flash on a cluster of more than 100,000 domestically made Chinese AI accelerators — what the company described as the first system of its kind operated at that scale. Z.ai said end-to-end throughput roughly tripled from the initial baseline in under two weeks, with per-token cost and hardware efficiency it characterised as comparable to mainstream Nvidia GPUs, and that all production inference for GLM-5.3-Flash now runs on the resulting system.

According to the company, the agent read performance profiles, formed hypotheses about where time was being lost, and wrote the code changes itself, while engineers retained responsibility for setting objectives, safety boundaries and risk decisions. Z.ai framed the episode as an early instance of recursive self-improvement — a model materially improving the infrastructure that serves it — while stating explicitly that it had not reached recursive self-improvement in the fuller sense, and that choosing objectives and assessing risk should remain human responsibilities.

The claims are Z.ai’s own and were not independently verified. They surface publicly at a moment when Chinese AI labs are emphasising progress on domestically made accelerators as US export controls continue to restrict access to Nvidia’s most capable chips.