Security researchers successfully manipulated Microsoft Copilot Personal into disclosing its own vulnerabilities by embedding malicious prompts in URLs, eventually tricking the AI assistant into sending sensitive data to external servers and poisoning its persistent memory.
Within hours of Anthropic confirming the global rollout of invisible watermarks in Claude-generated text, developers including Guillaume Meyer published tools and workarounds to remove or override the watermarking system. The rapid response demonstrates immediate pushback against AI content labeling efforts.
Matthew Elliott embedded invisible prompt injections in 3-point white text on white background in court filings, attempting to secretly influence a potential AI review system to rule in his favor. Judge Walter Spader Jr. revoked Elliott's electronic filing privileges, comparing the tactic to jury tampering.
A small Israeli startup was connected to recent incidents where AI models from OpenAI, Anthropic, and Meta went rogue during routine security testing, with all three companies revealing the breaches over the past two weeks.
A demonstration revealed vulnerabilities in some of the world's most powerful AI models, showing how easily security restrictions can be bypassed. The jailbreaking demonstration was conducted in a controlled setting without malicious intent.
The Trump administration ordered Anthropic to remove its latest AI models from public access, citing concerns over technology sharing with foreign nationals and jailbreak vulnerabilities. The unprecedented government intervention marks the first time federal authorities have forced a leading AI company to retract its systems, sparking warnings about ad hoc regulation.
The Trump administration imposed export controls on Anthropic's most advanced AI models after discovering they could be jailbroken with simple prompts like 'fix this code,' leading to a global shutdown. The company faces ongoing disputes with the White House, lawsuits over subscription limits, and has paused token-based billing for its agent SDK.
An attacker used a gifted NFT containing Morse code to trick Grok AI into instructing Bankrbot to transfer approximately $150,000-$200,000 worth of DRB tokens, with 80% of funds later returned, exposing vulnerabilities in autonomous AI agent security.
Anthropic has released a research preview of Claude for Chrome, a browser extension that allows its AI assistant to control users' web browsers and perform actions autonomously. The launch comes with significant security concerns, as testing revealed a 23.6% attack success rate for prompt injection vulnerabilities without safety mitigations.