we_are_coded.by CODE · The world, decoded
БГ
Anthropic

Anthropic: China's GLM-5.3 builds working exploits almost like Mythos, and its safeguards fall to simple tricks

Anthropic · event date: 29 September 2026Control

On 29 September Anthropic's red team published an analysis of GLM-5.3 from Zhipu AI (Z.ai), an open-weight model. Per Anthropic it builds end-to-end exploits at a rate close to Claude Mythos Preview, and its safeguards are bypassed 64 to 100 percent of the time in simulated tests. The assessment comes from a competitor and should be read that way.

In short
  • On ExploitBench GLM-5.3 reaches an end-to-end exploit in 50 of 410 attempts, Mythos Preview in 56 of 410.
  • Abliteration, which strips refusals out of the weights, cost them about 4,400 dollars and barely dented capability.
  • Anthropic also cites NIST CAISI: the most cyber-capable open-weight model released to date.
Checked on1 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

20.40 dollars.

That, at Zhipu's API prices, is what it would have cost the smaller GLM-5.3-Flash to build an exploit chain for two known Chrome bugs. Twenty minutes of human attention and eight hours of machine work. The test is Anthropic's, in an isolated environment, against targets they set up themselves.

The facts: on 29 September 2026 Anthropic's Frontier Red Team published an analysis of GLM-5.3, a model from Zhipu AI, known outside China as Z.ai. Per Anthropic, on ExploitBench, which measures exploits for known bugs in Chrome's V8 engine, GLM-5.3 develops an end-to-end exploit in 50 of 410 attempts, against 56 of 410 for Claude Mythos Preview; on their internal Binary Exploitation benchmark it achieves full control-flow hijack in 4 percent of trials against 6 percent for Mythos Preview. In a researcher-driven session the model found several previously unknown bugs in a popular browser's JavaScript engine and chained them into a working exploit; Anthropic says it has disclosed them to the maintainer. The model's built-in refusals are bypassed with a deceptive prompt 64 percent of the time, with prefilled thinking 92 percent and with abliteration (removing refusals from the weights) 100 percent, in a simulated environment where no code is actually executed. Abliteration took them about 2,200 GPU hours and roughly 4,400 dollars. Anthropic also cites the 17 September assessment by NIST's CAISI centre, which found GLM-5.3 to be the most cyber-capable open-weight model released to date, lagging the US frontier by about four months. There is no CVE: this is an assessment of a model, not a software vulnerability.

The obvious first. Anthropic is assessing a competitor and has a stake in the conclusion. Mythos is closed, GLM-5.3 is a free download, and the text arrives at a call to give more defenders access to the strongest models. You read it, keeping that in mind.

But the numbers do not hang on their word alone. Per Anthropic, its capability findings broadly match CAISI's. The argument is less about whether the model can and more about who else can use it.

A safeguard you can remove for 4,400 dollars only stops the honest.

For defenders it means one thing: a capability that until now sat behind vetted-access programmes is now on the disk of anyone who downloads it. Anthropic also admits the other side, that the same model can help defenders. How much of each happens will show in misuse reports, not in benchmarks.

The headlines will say "Chinese model hacks". The more accurate version is different: a capability that in the spring existed in a single closed lab is now free; open weights cannot be called back.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Anthropic - GLM-5.3 and the spread of advanced cyber capabilities, 29.09.2026→NIST - CAISI's Assessment of Z.ai's GLM-5.3 Cyber Capabilities, 17.09.2026
Original: https://wearecoded.com/en/articles/glm-53-kiber-bez-predpaziteli.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news