An analysis by the UK AI Security Institute and other researchers of 192 language models across 5,000 questions reveals major flaws in AI safety benchmarks. The study shows that a single safety score hides independent traits, and fewer than 2% of test questions actually differentiate models. To address these issues, researchers proposed adaptive testing to cut costs by up to 99% and a 'person-fit' check that detects simulated model 'sandbagging' with up to 97% accuracy, highlighting the need for more rigorous, psychological-grade testing standards.
Aug 27, 2026 · 5 sources
Aug 27, 2026 · 6 sources
Aug 27, 2026 · 4 sources
Aug 26, 2026 · 3 sources
Story comments
Loading comments…