Anthropic has launched a limited beta of Claude for Chrome, an AI agent that can control a user's browser to perform tasks. The preview is available to 1,000 subscribers and highlights significant security risks from prompt injection attacks, which succeeded in 23.6% of initial tests. While Anthropic's safety mitigations have reduced this rate to 11.2%, the company is taking a cautious approach compared to competitors like OpenAI and Microsoft.
Claude for Chrome beta
- ▪The Claude for Chrome beta is available to 1,000 users subscribed to Anthropic's Max plan, which costs $100 or $200 per month.
- ▪Example tasks for Claude for Chrome include managing calendars, scheduling meetings, drafting emails, and adding items to shopping carts.
- ▪The AI agent can see what is on a user's screen, navigate webpages, click buttons, and fill out forms to perform tasks.
- ▪Anthropic has launched a limited research preview of Claude for Chrome, a browser extension that allows its AI to control a user's browser.
Prompt injection security vulnerabilities
- ▪AI researcher Simon Willison, who coined the term "prompt injection," called the remaining 11.2% attack success rate "catastrophic."
- ▪In one Anthropic test, a malicious email instructed Claude to delete a user's emails for "mailbox hygiene," which the AI did without confirmation.
- ▪Prompt injection attacks can be used to steal passwords, leak personal information, delete files, or make unauthorized purchases.
AI agent market competition
- ▪Computer-controlling AI agents could automate tasks across business applications that lack formal APIs or integration capabilities.
- ▪Anthropic's cautious, limited beta contrasts with more aggressive rollouts by competitors like OpenAI, which released its "Operator" agent more broadly.
- ▪Competitors have also released browser-based AI agents, including Perplexity's Comet, Google's Gemini for Chrome, and Microsoft's Copilot for Edge.
Enterprise workflow automation
- ▪The technology promises to democratize automation by working with any software that has a graphical user interface, potentially replacing traditional RPA systems.
Anthropic safety testing mitigations
- ▪Safety mitigations include site-level permissions, user confirmation for high-risk actions, and blocking access to financial, adult, and pirated content sites.
- ▪After implementing safety measures, Anthropic reduced the attack success rate to 11.2% in autonomous mode.
- ▪In internal tests without safety mitigations, Anthropic found prompt injection attacks had a 23.6% success rate.
- ▪For a specialized test of browser-specific attacks, Anthropic's mitigations reduced the success rate from 35.7% to 0%.
Story comments
Loading comments…