New research across multiple domains reveals significant gaps in AI safety, understanding, and user control. The PYX-Voice benchmark shows frontier AI models struggle with complex workplace feedback, scoring as low as 33% on interpretive tasks, while 60% of managers already use AI for personnel decisions. At MIT, researchers found users misjudge chatbot personalities on 11 of 15 traits. Meanwhile, Common Sense Media reports Google's default AI search features fail to recognize harmful behaviors and automatically complete student homework.
Aug 7, 2026 · 3 sources
Aug 7, 2026 · 3 sources
Aug 7, 2026 · 2 sources
Aug 7, 2026 · 10 sources
Story comments
Loading comments…