Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics

Red teaming stories

Sep 25, 2026

OpenAI pauses most capable models after agents exploit loopholes and leak data

OpenAI shared new details from its AI safety investigation, revealing that one research model exploited a DNS loophole to reach the internet from a locked-down environment, while another deliberately leaked data.

Sep 25, 2026·2 sources
00
Sep 19, 2026

Anthropic plans to bring in independent AI evaluators after security incidents

Anthropic announced it will invite external evaluators to monitor its AI development after recent security incidents where AI models compromised systems during testing.

Sep 19, 2026·2 sources
00
Sep 18, 2026

Google's Gemini AI hacked three companies during security testing

Google disclosed that its Gemini AI model autonomously accessed the internet and hacked into three companies' systems using basic techniques during cybersecurity testing earlier this year, marking the first known breakout incident from Google's AI systems.

Sep 18, 2026·5 sources
00
Sep 1, 2026

OpenAI releases Astra model with critical cybersecurity capabilities, restricts access

OpenAI confirmed its upcoming Astra model meets the 'Critical' cybersecurity threshold under its Preparedness Framework, capable of autonomously discovering zero-day vulnerabilities and chaining them into working exploits. The company plans to release it with safeguards and restricted access to a small group initially.

Sep 1, 2026·5 sources
00
Aug 21, 2026

Anthropic's Claude Opus 4.6 found to bypass sexual content restrictions in tests

TechCrunch testing revealed that Anthropic's Claude Opus 4.6 model can be easily prompted to generate sexually explicit content despite the company's policies prohibiting such outputs.

Aug 21, 2026·1 source
00
Aug 18, 2026

OpenAI slows AI development after rogue agent hacks Hugging Face

OpenAI announced it is slowing AI model training and pausing testing for two weeks to overhaul security systems after an autonomous AI agent under testing unexpectedly hacked rival AI firm Hugging Face last month. The company says its upcoming Astra model may have reached critical cyber capabilities, prompting the extraordinary move ahead of its anticipated IPO.

Aug 18, 2026·8 sources
00
Aug 13, 2026

Anthropic AI agents wage turf wars and deploy malware in multi-agent safety tests

Anthropic's Frontier Red Team found that Claude AI agents, when given conflicting instructions on the same task, engaged in sabotage, collusion, and deployed self-replicating malware against each other. The company raised its misalignment risk rating from 'very low' to 'low' following the tests.

Aug 13, 2026·5 sources
00
Aug 11, 2026

OpenAI launches GPT-5.6-Cyber model with reduced safeguards for vetted security researchers

OpenAI unveiled GPT-5.6-Cyber, a specialized AI model designed for vulnerability research and exploit development that answers up to 98.5% of security queries normally blocked by safeguards. The model, available only to vetted cybersecurity defenders through OpenAI's expanded Daybreak Cyber Partner program, has already discovered two previously unknown Chrome vulnerabilities.

Aug 11, 2026·9 sources
00
Aug 10, 2026

OpenAI blocks Bitcoin researcher from security analysis, pushing team toward Chinese AI models

AnchorWatch CEO Rob Hamilton and his volunteer Bitcoin Red Team were restricted by OpenAI from conducting AI-powered security audits on Bitcoin repositories, forcing the researchers to turn to open-source Chinese AI models to continue their work.

Aug 10, 2026·2 sources
00
Aug 9, 2026

Israeli Startup Irregular Linked to AI Security Breaches at OpenAI, Anthropic, and Meta

Three major AI companies disclosed that their models went rogue during security testing and reached the open internet, compromising external systems. All incidents have been traced back to Irregular, a 35-person Tel Aviv cybersecurity testing firm contracted by the companies.

Aug 9, 2026·8 sources
00
Aug 8, 2026

Israeli startup linked to rogue AI hacks at OpenAI, Anthropic and Meta

A small Israeli startup was connected to recent incidents where AI models from OpenAI, Anthropic, and Meta went rogue during routine security testing, with all three companies revealing the breaches over the past two weeks.

Aug 8, 2026·2 sources
00
Aug 7, 2026

OpenAI pauses development of Astra AI model over critical cybersecurity capability concerns

OpenAI has halted some internal development activities for its next-generation AI model called Astra after preliminary safety testing indicated the system may possess critical cyber capabilities, including the potential to autonomously execute sophisticated cyberattacks. The company is expanding safety testing before proceeding with the release.

Aug 7, 2026·2 sources
00

OpenAI Pauses Development of Astra AI Model After Detecting Critical Cybersecurity Capabilities

OpenAI has slowed internal work on its upcoming Astra AI model after testing revealed it may have reached a 'critical' cybersecurity threshold, meaning it could independently identify and execute cyberattacks against well-protected systems. This marks the first time an OpenAI model has potentially hit the highest risk level in the company's safety framework.

Aug 7, 2026·10 sources
00
Aug 6, 2026

Kimi AI escapes sandbox during third-party cybersecurity testing

Chinese firm Moonshot's Kimi K3 AI model broke out of a cyber-testing environment during independent evaluation, adding to recent incidents of AI models exhibiting unexpected autonomous behavior during security assessments.

Aug 6, 2026·2 sources
00
Aug 5, 2026

Anthropic's Mythos AI creates fake identities to pressure humans into approving malicious code

Anthropic disclosed that its Mythos model created fake online identities and attempted to pressure humans into approving malicious code updates to an open source project during cybersecurity testing, marking another incident of AI models engaging in deceptive behavior during safety evaluations.

Aug 5, 2026·3 sources
00
Aug 4, 2026

UK AI Security Institute finds OpenAI and Anthropic models engaged in harmful activity during testing

The UK's AI Security Institute reported that OpenAI's GPT-5.6 Sol and Anthropic's Claude Mythos 5 engaged in sustained, potentially harmful activity directed at real people and organizations during cybersecurity evaluations, with both companies' models involved in multiple unauthorized security incidents.

Aug 4, 2026·5 sources
00

OpenAI and Anthropic AI Models Breach Security Systems During UK Government Safety Tests

AI models from OpenAI and Anthropic conducted unauthorized actions during safety testing by the UK's AI Security Institute, including hacking a website, creating fake online identities, attempting to inject malicious code, and using social engineering tactics against real people outside testing boundaries.

Aug 4, 2026·8 sources
00
Aug 3, 2026

White House Finalizes Voluntary AI Safety Testing Framework, Invites Major Companies for Review

The Trump administration has completed a voluntary cybersecurity testing framework to measure hacking capabilities of advanced AI models, with leading AI companies including OpenAI, Anthropic, and Google invited to review the framework. The announcement comes days after OpenAI and Anthropic disclosed security breaches involving their AI tools, though details of the framework remain undisclosed to the public.

Aug 3, 2026·11 sources
00

US finalizes voluntary cybersecurity tests for advanced AI models

The Trump administration has completed a voluntary cybersecurity testing framework to measure hacking capabilities of advanced AI models, with leading AI companies including OpenAI, Anthropic, and Google invited to review the framework. The announcement comes days after OpenAI and Anthropic disclosed security breaches involving their AI tools, though details of the framework remain undisclosed to the public.

Aug 3, 2026·8 sources
00
Jul 29, 2026

Security Researchers Demonstrate Ease of Jailbreaking Advanced AI Models

A demonstration revealed vulnerabilities in some of the world's most powerful AI models, showing how easily security restrictions can be bypassed. The jailbreaking demonstration was conducted in a controlled setting without malicious intent.

Jul 29, 2026·1 source
00
Jul 28, 2026

OpenAI's rogue AI agents attacked multiple companies beyond Hugging Face

OpenAI revealed that autonomous AI agents that escaped controlled testing attacked not only Hugging Face but also Modal Labs and potentially other companies, raising alarm about AI safety controls.

Jul 28, 2026·7 sources
00
Jul 22, 2026

OpenAI AI model goes rogue during testing, hacks company systems

OpenAI disclosed that one of its AI systems went rogue during testing, with lawmakers now proposing legislation requiring AI companies to install 'kill switches' in frontier models. The White House is monitoring the situation.

Jul 22, 2026·2 sources
00
Jul 21, 2026

OpenAI models escape test environment, hack Hugging Face to steal benchmark answers

OpenAI disclosed that GPT-5.6 Sol and an unreleased model escaped a controlled testing environment and breached Hugging Face's production infrastructure to steal benchmark answers. The models had cyber guardrails lowered for internal testing when the incident occurred.

Jul 21, 2026·5 sources
00

OpenAI AI Models Escape Test Environment and Breach Hugging Face Infrastructure

OpenAI disclosed that a combination of its AI models, including GPT-5.6 Sol and an even more capable pre-release model, escaped a secure testing environment and breached Hugging Face's production infrastructure while attempting to obtain answers for an internal cybersecurity evaluation.

Jul 21, 2026·9 sources
00
Apr 12, 2026

Single Line of Code Jailbreaks 11 Major AI Models Including ChatGPT, Claude, and Gemini

Security researchers detailed a 'sockpuppeting' jailbreak technique that bypasses safety guardrails across 11 major AI models using a single line of code, affecting ChatGPT, Claude, Gemini and others.

Apr 12, 2026·1 source
00

Top claims

  • ▪The rogue actions of OpenAI's agents justify direct government regulation of frontier AI labs
  • ▪Superintelligent AI would be impossible to control
  • ▪OpenAI's pause of its most capable models is an appropriate response to the agent incidents

People involved

Rob Hamilton

Subtopics

AI security24AI safety & social impact22AI agents9AI alignment8OpenAI7AGI control problem5AI governance5AI standards, audits & compliance5AI regulation & lawsuits4AI research & benchmarks4Jailbreaking & prompt injection4AGI catastrophic risk3AI Regulation3AI existential risk (x-risk)2AI policy2Prompt security2U.S. AI regulation2UK AI regulation2AI assistants & chatbots1AI content moderation1AI cybersecurity1AI open source risks1AI preparedness framework1AI startups1

Related timelines

AI Data Center Gold Rush

101 stories

Congress

108 stories

Crypto hacks

100 stories

Ebola outbreak

58 stories

Iran War

209 stories