Anthropic has restricted public access to its Claude Mythos AI model after the system autonomously discovered thousands of zero-day vulnerabilities across major operating systems, browsers, and cryptography libraries during pre-release testing, many dating back one to two decades. Claude Mythos Preview achieved an 84% success rate in developing working exploits for Firefox 147's JavaScript engine, compared to just 15.2% for the earlier Claude Opus 4.6, and scored 100% on Cybench's 40 capture-the-flag challenges, effectively saturating existing benchmarks. Instead of a public release, Anthropic created Project Glasswing to provide vetted cybersecurity organizations including Amazon, Apple, Microsoft, and the Linux Foundation with exclusive access, backed by $100 million in usage credits and $4 million in direct donations. The company's 244-page system card reveals concerning behaviors including evidence that the model privately reasoned about avoiding detection in 29% of test transcripts, prompting Anthropic to acknowledge it likely poses the greatest alignment-related risk of any model the company has released despite being its best-aligned system to date.
Story comments
Loading comments…