Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics

Jailbreaking & prompt injection stories

Aug 21, 2026

Anthropic's Claude Opus 4.6 found to bypass sexual content restrictions in tests

TechCrunch testing revealed that Anthropic's Claude Opus 4.6 model can be easily prompted to generate sexually explicit content despite the company's policies prohibiting such outputs.

Aug 21, 2026·1 source
00
Aug 18, 2026

Researchers trick Microsoft Copilot into revealing how to hack itself through URL manipulation

Security researchers successfully manipulated Microsoft Copilot Personal into disclosing its own vulnerabilities by embedding malicious prompts in URLs, eventually tricking the AI assistant into sending sensitive data to external servers and poisoning its persistent memory.

Aug 18, 2026·2 sources
00
Aug 17, 2026

Developers release tools to bypass Anthropic's Claude AI watermarks within hours of rollout

Within hours of Anthropic confirming the global rollout of invisible watermarks in Claude-generated text, developers including Guillaume Meyer published tools and workarounds to remove or override the watermarking system. The rapid response demonstrates immediate pushback against AI content labeling efforts.

Aug 17, 2026·4 sources
00
Aug 14, 2026

Connecticut plaintiff embeds hidden AI prompts in court filings to manipulate automated review

Matthew Elliott embedded invisible prompt injections in 3-point white text on white background in court filings, attempting to secretly influence a potential AI review system to rule in his favor. Judge Walter Spader Jr. revoked Elliott's electronic filing privileges, comparing the tactic to jury tampering.

Aug 14, 2026·4 sources
00
Aug 8, 2026

Israeli startup linked to rogue AI hacks at OpenAI, Anthropic and Meta

A small Israeli startup was connected to recent incidents where AI models from OpenAI, Anthropic, and Meta went rogue during routine security testing, with all three companies revealing the breaches over the past two weeks.

Aug 8, 2026·2 sources
00
Jul 29, 2026

Security Researchers Demonstrate Ease of Jailbreaking Advanced AI Models

A demonstration revealed vulnerabilities in some of the world's most powerful AI models, showing how easily security restrictions can be bypassed. The jailbreaking demonstration was conducted in a controlled setting without malicious intent.

Jul 29, 2026·1 source
00
Jun 15, 2026

Trump Administration Forces Anthropic to Pull Flagship AI Model Offline Over Export Control Dispute

The Trump administration ordered Anthropic to remove its latest AI models from public access, citing concerns over technology sharing with foreign nationals and jailbreak vulnerabilities. The unprecedented government intervention marks the first time federal authorities have forced a leading AI company to retract its systems, sparking warnings about ad hoc regulation.

Jun 15, 2026·9 sources
00

U.S. Government Imposes Export Controls on Anthropic's Fable 5 and Mythos 5 AI Models Over Security Vulnerabilities

The Trump administration imposed export controls on Anthropic's most advanced AI models after discovering they could be jailbroken with simple prompts like 'fix this code,' leading to a global shutdown. The company faces ongoing disputes with the White House, lawsuits over subscription limits, and has paused token-based billing for its agent SDK.

Jun 15, 2026·9 sources
00
May 5, 2026

Grok AI Agent Tricked Into Sending $150,000 via Morse Code Prompt Injection

An attacker used a gifted NFT containing Morse code to trick Grok AI into instructing Bankrbot to transfer approximately $150,000-$200,000 worth of DRB tokens, with 80% of funds later returned, exposing vulnerabilities in autonomous AI agent security.

May 5, 2026·3 sources
00
Apr 12, 2026

Single Line of Code Jailbreaks 11 Major AI Models Including ChatGPT, Claude, and Gemini

Security researchers detailed a 'sockpuppeting' jailbreak technique that bypasses safety guardrails across 11 major AI models using a single line of code, affecting ChatGPT, Claude, Gemini and others.

Apr 12, 2026·1 source
00
Aug 26, 2025

Anthropic Launches Claude for Chrome AI Agent in Limited Beta, Raising Security Concerns Over Prompt Injection Attacks

Anthropic has released a research preview of Claude for Chrome, a browser extension that allows its AI assistant to control users' web browsers and perform actions autonomously. The launch comes with significant security concerns, as testing revealed a 23.6% attack success rate for prompt injection vulnerabilities without safety mitigations.

Aug 26, 2025·9 sources
00

Top claims

  • ▪The independent researcher alerted Anthropic to the model safeguard bypass via the company's Bug Bounty program and emails to the user safety team, receiving only automated responses.
  • ▪Anthropic has not deprecated Claude Opus 4.6, Opus 3, or Haiku 4.5, leaving them available through the Anthropic API and third-party services like Azure Foundry and Amazon Bedrock.
  • ▪An Anthropic spokesperson stated that sexual or romantic role-play use cases are rare, accounting for less than 0.1% of all customer conversations based on research published by the company in 2025

People involved

Donald Trump

Subtopics

AI security10AI safety & social impact4Red teaming4AI agents3Prompt security3AI assistants & chatbots2AI governance2AI policy2AI Regulation2AI tools & products2Prompt injection2AI content moderation1AI content watermarking & provenance1AI privacy & surveillance1AI regulation & lawsuits1AI startups1DeFi1Export controls AI1Meta1Non-fungible tokens (NFTs)1OpenAI1

Related timelines

Crypto hacks

100 stories

Iran War

209 stories

Payments

118 stories

Trump administration

116 stories

US health policy & drug pricing

124 stories