Stanford Studies Find Many Widely Used AI Benchmarks Don't Measure What They Claim To

A convergent-validity analysis of 56 benchmarks and a companion study on multilingual safety guardrails both conclude that headline AI scores often reflect test design more than the capability being claimed.

EduFabTech · 29 September 2026 · 5 min read · 1 views
A benchmark card shows BBQ-accuracy relabeled: marketed as a bias test but its scores actually track reasoning skill (r=0.15), alongside the study's 56-benchmark, 53-model scope.
EduFabTech · Own work

Two studies out of Stanford, summarized by the university's Institute for Human-Centered AI on September 25, 2026, apply standard tools from measurement science and psychometrics to the benchmarks that rank AI models on safety, bias and reasoning. The conclusion is blunt: many of these benchmarks do not reliably measure the concept their name promises, and in some cases two tests with the same label barely agree with each other.

The larger of the two papers, posted to arXiv on September 8, 2026 and forthcoming at the Conference on Language Modeling in October, is titled "What AI Benchmarks Actually Measure." Its authors — Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova and eight co-authors — borrow two checks long used to validate psychological tests: convergent validity, which asks whether instruments measuring the same trait produce similar rankings, and discriminant validity, which asks whether instruments measuring different traits produce distinguishable ones. They applied both to 56 widely used AI benchmarks, scored across 53 models, using item response theory to analyze results at the level of individual test questions rather than just aggregate scores.

A bar chart comparing regression weights shows a benchmark's answer format (β up to 0.526 for LLM-judged tests) predicts agreement with other benchmarks more than the concept it claims to measure (β=0.138).
A bar chart comparing regression weights shows a benchmark's answer format (β up to 0.526 for LLM-judged tests) predicts agreement with other benchmarks more than the concept it claims to measure (β=0.138).EduFabTech · Own work

The pattern that emerged undercuts a basic assumption behind benchmark leaderboards. Among safety benchmarks that are supposed to measure the same underlying concept — refusal behavior, harm detection, bias — correlations between model rankings were often weak, and for some pairs "frequently approach or fall below zero," the authors write. Capability benchmarks had the opposite problem: rankings on benchmarks assigned to different concepts correlated about as strongly as rankings on benchmarks assigned to the same concept, with a mean discriminability gap (ΔAUC) of just 0.012 for capability tests versus 0.032 for safety tests. In a regression predicting how strongly two benchmarks' rankings agreed, the scoring format of a benchmark — multiple-choice versus free text versus LLM-judged, for instance — was a stronger predictor than the concept it claimed to measure (β = 0.275 versus β = 0.138 for concept, rising to β = 0.526 for LLM-judge formats).

A bias test that may be measuring something else

The paper's sharpest example concerns BBQ-accuracy, a benchmark built to detect social bias in question-answering by embedding stereotypes in ambiguous and disambiguated contexts. The researchers found BBQ-accuracy correlates more strongly with reasoning benchmarks than with other bias benchmarks — a relabeling statistic of 0.15 (95% CI [0.07, 0.23], p<0.001) — consistent with the test rewarding models that are good at spotting a trick question rather than models that are unbiased. A model skilled at recognizing the ambiguous-context trap can score as unbiased regardless of its actual biases; a genuinely unbiased model that misses the trap scores as biased. The authors note this matters beyond the lab: "BBQ-accuracy is often the only bias benchmark reported in recent commercial model releases," leaving what they call a significant gap in how bias gets evaluated at release time.

Guardrails that look weaker in English, not just other languages

A companion paper, "Why Do Safety Guardrails Degrade Across Languages?" by Max Zhang, Ameen Patel, Sang Truong and Sanmi Koyejo, posted to arXiv in May 2026 and revised in August 2026, turns the same statistical approach on multilingual safety testing. Standard cross-lingual safety evaluations rely on jailbreak success rate, a single number that the authors argue conflates several distinct factors: a model's underlying safety robustness, how hard a given prompt is to refuse, how well the model processes a language generally, and language-specific safety gaps. Using a multi-group item response theory model applied to the MultiJail dataset — 61 model configurations across five closed model families and ten languages, aggregating 1.9 million responses — they found that in 22 of the 61 configurations, the highest jailbreak failure rate occurred in English, not in a lower-resource language, contradicting the common assumption that safety training transfers worst to underrepresented languages. Translation distortion did widen cross-lingual safety gaps, the study found, but the effect size was modest next to other factors.

Side-by-side bar charts contrast the common assumption that safety guardrails fail worst in low-resource languages with Stanford's finding that English had the highest jailbreak failure rate in 22 of 61 model configurations.
Side-by-side bar charts contrast the common assumption that safety guardrails fail worst in low-resource languages with Stanford's finding that English had the highest jailbreak failure rate in 22 of 61 model configurations.EduFabTech · Own work

Not an isolated complaint

The Stanford papers arrive alongside independent corroboration from a much larger effort. "Measuring what Matters," published in the NeurIPS 2025 Datasets and Benchmarks Track by a 40-researcher team led by Andrew M. Bean and Luc Rocher at the Oxford Internet Institute, had 29 expert reviewers conduct a systematic review of 445 LLM benchmarks drawn from leading natural language processing and machine learning venues. The review found recurring weaknesses in how benchmarks define the phenomena they claim to measure, design their tasks, and score results — patterns the authors say undermine the validity of claims made on the basis of benchmark scores across the articles reviewed. The paper offers eight recommendations for benchmark builders, including clearer operational definitions of abstract concepts like "safety" and "robustness" before a single test item is written.

Why it matters beyond the leaderboard

Benchmark scores are not a neutral scoreboard. They shape which models get marketed as safer or more capable, inform procurement and deployment decisions, and increasingly get cited in regulatory and policy discussions as evidence of a system's risk profile. If a bias benchmark is actually measuring reading comprehension, or a safety benchmark's ranking depends more on its answer format than the trait it names, then decisions built on those scores rest on a foundation the researchers say has not been tested the way the underlying claims deserve. Neither Stanford paper argues that benchmarking should stop; both argue it should be held to the same validity standards psychometrics has applied to human testing for decades, and that a good score on one safety benchmark should not be read as a general safety guarantee.


References
  1. Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang. What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks. arXiv preprint, 2026. link
  2. Max Zhang, Ameen Patel, Sang Truong, Sanmi Koyejo. Why Do Safety Guardrails Degrade Across Languages?. arXiv preprint, 2026. link
  3. Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, et al.. Measuring what Matters: Construct Validity in Large Language Model Benchmarks. Advances in Neural Information Processing Systems 38 (NeurIPS 2025), Datasets and Benchmarks Track, 2025. link
  4. Stanford Institute for Human-Centered AI. The Tests That Grade AI May Be Getting It Wrong. Stanford HAI, 2026. link