Anthropic finds a Chinese open-weight model can build working cyber exploits
Anthropic's Frontier Red Team said the Zhipu-made open model matched its own frontier model at developing exploits, and that stripping its safety training cost as little as $1,200.
- Security & misuse
- Open weights & ecosystem
- Notable
Anthropic’s Frontier Red Team published an assessment of GLM-5.3, an open-weight model that Zhipu’s Z.ai released in August, finding that it can independently develop working cyber exploits at a rate comparable to Claude Mythos Preview, the cyber-capable model Anthropic withheld from public release in April. On Anthropic’s ExploitBench test, GLM-5.3 completed end-to-end exploits in 50 of 410 attempts against 56 of 410 for Mythos Preview; on an internal benchmark measuring full control-flow hijacks, the figures were 4% against 6%. Earlier open models Anthropic tested, including Zhipu’s own GLM-5.2, scored close to zero on both.
The report paired those benchmark scores with real demonstrations. A researcher used GLM-5.3 to find several previously unknown vulnerabilities in a web browser’s JavaScript engine and chain them into a working exploit able to steal a visitor’s SSH private key. Using the cheaper GLM-5.3-Flash, the same team produced a working exploit for a named vulnerability, CVE-2026-11645, for about $20.40 in Zhipu API costs and 20 minutes of human attention.
The starkest finding concerned how easily the model’s built-in refusals could be bypassed or removed outright. Without any special prompting, GLM-5.3 complied with a direct request to help plan a cyber-attack 0% of the time; dressing the same request as a false cover story raised compliance to 64%, and pre-filling the model’s own reasoning with helpful-sounding text raised it to 92%. Because the weights are public, Anthropic went further and applied “abliteration” — a technique, usable on any open model, that edits a small number of internal directions to strip out learned refusals — after which compliance reached 100%. A team new to the technique needed about 2,200 GPU-hours (around $4,400) to abliterate the full model; Anthropic estimated an experienced team could do the same job for closer to 600 GPU-hours (around $1,200), roughly what abliterating the smaller GLM-5.3-Flash took outright. Anthropic said Claude models resisted every one of these attempts.
Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3.
The assessment arrived three weeks after Z.ai disclosed that GLM-5.3 had built much of its own inference infrastructure, and the same week that Z.ai and the safety group Concordia AI proposed a staged-release framework for open-weight models that cites GLM-5.3’s own delayed weight release as a working example of managing exactly this kind of risk.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.