Daily desk

News

Artificial intelligence, technology, science, health and education β€” reported daily, with the original source attached to every item.

1 story

A benchmark card shows BBQ-accuracy relabeled: marketed as a bias test but its scores actually track reasoning skill (r=0.15), alongside the study's 56-benchmark, 53-model scope.
Latest πŸ€– Artificial Intelligence

Stanford Studies Find Many Widely Used AI Benchmarks Don't Measure What They Claim To

Applying psychometric validity tests to 56 AI benchmarks, Stanford researchers found safety-benchmark rankings barely correlate with each other, and a bias test used in commercial model releases seems to measure reading comprehension instead.

EduFabTech Β· 1h ago Β· 5 min read