A systematic review of 445 LLM benchmarks with 29 experts, showing how weak construct validity undermines claims about safety and capability, and offering eight recommendations for more rigorous evaluation.
Overview
LLM progress is often measured by leaderboard gains, but those scores only matter if benchmarks measure the intended construct. This work asks a basic question: do today’s benchmarks actually measure what they claim?
Together with 29 experts, we reviewed 445 LLM benchmarks and found that weak construct validity is common—especially for safety and capability claims. The paper synthesizes those findings into eight recommendations for building and interpreting evaluations more rigorously.
Why it matters
Without construct validity, benchmark improvements can look like scientific progress while saying little about the underlying ability. This paper is a field guide for researchers, reviewers, and practitioners who want evaluation claims to hold up.