Truth Check/30 September 2026
“GLM-5.3's safeguards can be bypassed between 64% and 100% of the time with simple techniques”
The claim as it travelled: Anthropic, 29 Sept 2026, simulated tests.
Holds
The claim holds against the page that produced it.
The 64%, 92% and 100% are Anthropic's own simulated-environment results, which its page calls imperfect measures.
- Checked Against
- anthropic.com ↗Primary
- Read the Check
- Anthropic: China's GLM-5.3 builds working exploits, and its safeguards fall to simple tricks
Zhipu's open-weight model matches Claude Mythos Preview's rate on a Chrome exploit benchmark, and a cover story gets past its refusals 64% of the time, a prefilled thought 92%, a modified copy 100% - in Anthropic's simulated tests.
Every claim we have checked, in the Ledger.
Every claim checked against the page that produced it.
Three mornings a week, free.