OpenAI released its GPT-5.5-Cyber model (codenamed 'Spud') to vetted cybersecurity professionals on Thursday, matching capabilities of Anthropic's Mythos model released a month earlier. The AI Security Institute found GPT-5.5 achieved a 71.4% pass rate on expert-tier cybersecurity tasks compared to Mythos's 68.6%, with both models completing sophisticated 32-step corporate network simulations that would take human experts around 20 hours. Security testing revealed GPT-5.5 solved a 12-hour reverse-engineering puzzle in 10 minutes for $1.73 in API costs, but researchers also discovered a universal jailbreak bypassing all safety guardrails after six hours of red-teaming. The UK government responded with £90 million in new cyber resilience funding and guidance warning organizations to prepare for AI-accelerated vulnerability discovery, as 43% of UK businesses suffered cyber attacks in the past year.
GPT-5.5's autonomous cyberattack capabilities and benchmark performance
- ▪Anthropic's Claude Mythos Preview achieved a 68.6% pass rate on the AI Security Institute's Expert tier cybersecurity tasks
- ▪The AI Security Institute built the 32-step corporate network simulation with cybersecurity firm SpecterOps
- ▪The AI Security Institute's 32-step corporate network simulation requires chaining together reconnaissance, credential theft, lateral movement across multiple Active Directory forests, a supply-chain pivot through a CI/CD pipeline, and exfiltration of a protected internal database
- ▪OpenAI's GPT-5.5 completed a 32-step corporate network simulation in two out of 10 attempts
- ▪Anthropic's Claude Mythos Preview completed the AI Security Institute's 32-step corporate network simulation in three of 10 attempts
- ▪OpenAI's GPT-5.5 achieved a 71.4% average pass rate on the AI Security Institute's Expert tier cybersecurity tasks
- ▪OpenAI's GPT-5.5 cracked a reverse-engineering security puzzle in 10 minutes and 22 seconds that took a human security expert approximately 12 hours
- ▪OpenAI released the GPT-5.5-Cyber model to vetted security defenders
- ▪OpenAI's GPT-5.5 model can autonomously execute sophisticated cyberattacks
- ▪OpenAI's GPT-5.5 is roughly on par with Anthropic's Claude Mythos for offensive cyber capabilities according to the AI Security Institute
- ▪OpenAI's GPT-5.5 reverse-engineering solution cost $1.73 in API usage
- ▪The AI Security Institute estimates the 32-step corporate network simulation would take a human expert around 20 hours to complete
Safety concerns and jailbreak vulnerabilities
- ▪Researchers identified a universal jailbreak that bypassed OpenAI's GPT-5.5 safety guardrails entirely across all malicious cyber queries tested
UK cybersecurity landscape and government response
- ▪The UK government announced £90 million in new funding to boost cyber resilience
- ▪43% of UK businesses suffered a cyber breach or attack in the past 12 months according to the UK government's annual Cyber Security Breaches Survey
- ▪The AI Security Institute published findings on OpenAI's GPT-5.5 on Thursday
Perspective of UK government
- ▪The UK government announced £90 million in new funding to boost cyber resilience
Perspective of OpenAI
- ▪OpenAI released the GPT-5.5-Cyber model to vetted security defenders
Perspective of Cybersecurity researchers
- ▪Researchers identified a universal jailbreak that bypassed OpenAI's GPT-5.5 safety guardrails entirely across all malicious cyber queries tested
Story comments
Loading comments…