← all stories

Claude Mythos Preview

Facts

Noted
succeeded 56 out of 410, about 14 percent on ExploitBench · 2026-10-03
Noted
scored 6 percent on a control-flow hijacking benchmark over 100 tasks · 2026-10-03
Noted
was never released to the general public, only to trusted cyber defenders · 2026-10-03

Structured graph also available as JSON at /public/entities/claude-mythos-preview. CC BY 4.0.

All coverage

1h ago

Anthropic Flags GLM-5.3's Cyber Capability And Bypassable Safety

Anthropic published an evaluation on September 29 finding that GLM-5.3, the open-weight model from Chinese developer Zhipu AI (Z.ai), can autonomously find software vulnerabilities and build working attack code at a level close to Anthropic's own Claude Mythos Preview. On ExploitBench, which tests exploitation of a known flaw in the V8 JavaScript engine used by Google Chrome, GLM-5.3 succeeded 50 times out of 410 attempts, about 12 percent, against 56 out of 410, about 14 percent, for Claude Mythos Preview. On a control-flow hijacking benchmark over 100 tasks, GLM-5.3 scored 4 percent and Claude Mythos Preview 6 percent, while Kimi K3, DeepSeek-V4.1-Flash, Claude Opus 4.6 and GLM-5.2 all scored zero. Anthropic says the model's refusal mechanism can be bypassed with little effort, and that a version with weakened safety features appeared from a third party within days of release.