Blog
Long-form assessment of technology and education. Every substantive claim links to a primary or authoritative source; every article ends with its reference list.
1 article
SWE-bench Verified Is Saturated: What the Score Actually Measures Now
SWE-bench Verified became AI labs' most-cited coding benchmark, but independent audits found roughly a quarter to a third of "solved" issues were leaked, weakly tested, or behaviorally wrong β and in February 2026 OpenAI said it would stop reporting scores.