NeurIPS 2025 Datasets & Benchmarks

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Rose Kirk, Fangru Lin, Gabrielle Kaili-May Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yushi Yang, Yilun Zhao, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob Nicolaus Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip Torr, Cozmin Ududec, Luc Rocher, Adam Mahdi

A systematic review of 445 LLM benchmarks with 29 experts, showing how weak construct validity undermines claims about safety and capability, and offering eight recommendations for more rigorous evaluation.

Overview

LLM progress is often measured by leaderboard gains, but those scores only matter if benchmarks measure the intended construct. This work asks a basic question: do today’s benchmarks actually measure what they claim?

Together with 29 experts, we reviewed 445 LLM benchmarks and found that weak construct validity is common—especially for safety and capability claims. The paper synthesizes those findings into eight recommendations for building and interpreting evaluations more rigorously.

Why it matters

Without construct validity, benchmark improvements can look like scientific progress while saying little about the underlying ability. This paper is a field guide for researchers, reviewers, and practitioners who want evaluation claims to hold up.

Paper

Read the PDF on OpenReview