UK AI Security Institute finds major flaws in language model safety benchmarks
Researchers at the UK AI Security Institute used psychometric methods to demonstrate that popular safety benchmarks for language models don't measure one consistent trait, and that blanket blocking of requests can artificially inflate safety scores while reducing practical utility.