AIThe Decoder2h ago

Psychological methods reveal major weaknesses in AI security testing

Psychological methods reveal major weaknesses in AI security testing

Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for language models don't measure one consistent trait. Blanket blocking of requests can artificially inflate a safety score even as the model gets less useful day to…

Read full article

Source: The Decoder · Opens in new tab