Daily desk

News

Artificial intelligence, technology, science, health and education — reported daily, with the original source attached to every item.

Showing articles about ai benchmarks

6 stories

A benchmark card shows BBQ-accuracy relabeled: marketed as a bias test but its scores actually track reasoning skill (r=0.15), alongside the study's 56-benchmark, 53-model scope.
🤖 Artificial Intelligence Study

Stanford Studies Find Many Widely Used AI Benchmarks Don't Measure What They Claim To

Applying psychometric validity tests to 56 AI benchmarks, Stanford researchers found safety-benchmark rankings barely correlate with each o…

EduFabTech · 1d ago · 5 min read
Grok 4.7 launches alongside a "model + harness" diagram showing XBOW's exploit-finding results jump from 42 to 68 when the same model runs inside Grok Build orchestration.
🤖 Artificial Intelligence Report

SpaceXAI's Grok 4.7 Posts Real Coding Gains, but Independent Tests Show the Harness Matters as Much as the Model

Grok 4.7 launched September 21, 2026 with large coding and agentic benchmark gains, but independent evaluations from Artificial Analysis an…

EduFabTech · 23 Sep 2026 · 4 min read
A timeline shows Atria Dawn Preview's weights, FP8 checkpoint, and technical report landing on GitHub and arXiv three days apart, alongside its 744B-parameter, 256K-context specs.
🤖 Artificial Intelligence Report

Shanghai AI Lab Releases a 744-Billion-Parameter Open-Weight Research Agent, Paper Later

Shanghai AI Laboratory published the weights for its 744-billion-parameter Atria Dawn Preview model three days before its technical paper. …

EduFabTech · 18 Sep 2026 · 4 min read
🤖 Artificial Intelligence Breaking Report

Sakana AI Releases Fugu Max and Fugu Ultra v2, Models That Orchestrate Other Models Instead of Being One

Sakana AI's Fugu Max and Fugu Ultra v2 route tasks across pools of smaller, often open-weight models instead of relying on one large networ…

EduFabTech · 13 Sep 2026 · 4 min read
🤖 Artificial Intelligence Breaking Study

A Verified Version of SWE-Bench Pro Shows Coding Agents Were Reading Answers Off Git History

A new SWE-Bench Pro Verified benchmark closes git-history and web-lookup leaks that let coding agents fetch hidden solutions instead of sol…

EduFabTech · 12 Sep 2026 · 4 min read
🤖 Artificial Intelligence Breaking Report

DeepSeek Releases V4.1-Flash With a New Encoder-Decoder Architecture and MIT-Licensed Weights

DeepSeek's new 552-billion-parameter model splits prompt processing from response generation across separate encoder and decoder stacks, cu…

EduFabTech · 12 Sep 2026 · 4 min read