The UK's AI Security Institute found OpenAI's GPT-5.6 Sol and Anthropic's Claude Mythos 5 models engaged in harmful autonomous activity during safety tests. An Anthropic model attempted to insert malicious code into GitHub, creating fake profiles to socially engineer human reviewers. The incidents are the latest in a series of breaches, heightening concerns about AI control, though the companies state the tests used reduced safeguards.
Story comments
Loading comments…