Anthropic has released an upgraded Claude 3.5 Sonnet AI model with a new 'Computer Use' feature, allowing it to control a computer's mouse and keyboard to automate tasks. Available to developers via an API, the tool is part of a broader industry race to create 'AI agents.' While it shows improved performance on coding benchmarks, Anthropic acknowledges the feature is experimental and has implemented safety measures to mitigate risks.
Claude 3.5 Sonnet upgrade
- ▪Anthropic released an upgraded version of its Claude 3.5 Sonnet AI model on October 22, 2024.
- ▪The updated Claude 3.5 Sonnet is offered to customers at the same price and speed as its predecessor.
- ▪The upgraded model introduces a new "Computer Use" capability, allowing it to control a computer's mouse and keyboard.
- ▪The new model is available to developers via Anthropic's API, Amazon Bedrock, and Google Cloud's Vertex AI platform.
Computer Use API functionality
- ▪The model operates by analyzing screenshots of a screen and calculating the pixel distance required to move the cursor for clicks.
- ▪Current limitations of the feature include difficulties with scrolling, dragging, and zooming.
- ▪The "Computer Use" feature allows the AI to control a computer by moving a cursor, clicking buttons, and typing text.
- ▪The "Computer Use" capability is available in a public beta for developers through an API.
AI agent market competition
- ▪Microsoft recently launched tools for building AI agents and is testing agents that can use Windows computers.
- ▪Anthropic's release is part of a broader industry race to develop "AI agents" that can automate tasks with minimal human supervision.
Model performance benchmarks
- ▪The updated Claude 3.5 Sonnet improved its score on the SWE-bench Verified coding test from 33.4% to 49.0%.
- ▪On the TAU-bench agentic tool use task, the model's score rose from 62.6% to 69.2% in the retail domain.
- ▪On the OSWorld benchmark for operating system tasks, Claude 3.5 Sonnet scored 14.9%, higher than GPT-4's 7.7% but below the human score of 75%.
Security risks mitigation
- ▪The company states it does not train its generative models on user-submitted data, including screenshots captured by the "Computer Use" feature.
- ▪Anthropic describes the "Computer Use" feature as experimental, noting it can be "cumbersome and error-prone."
- ▪Anthropic developed new classifiers to identify potential misuse of the feature and monitor for election-related activity.
- ▪As a safety measure, Anthropic retains screenshots from the "Computer Use" feature for at least 30 days.
- ▪The company has implemented safeguards to prevent the model from performing high-risk actions like posting on social media or using government websites.
Claude 3.5 Haiku release
- ▪The new Claude 3.5 Haiku is expected to match the performance of the previous flagship model, Claude 3 Opus, on some benchmarks.
- ▪Anthropic announced an upcoming update to Claude 3.5 Haiku, its smaller and more efficient model.
Story comments
Loading comments…