Blog
Long-form assessment of technology and education. Every substantive claim links to a primary or authoritative source; every article ends with its reference list.
2 article(s)
Long-Context Windows: What Independent Benchmarks Show Versus What Vendors Claim
Context windows now advertise millions of tokens, but independent benchmarks — RULER, NoLiMa, Lost in the Middle and a newer preprint on mu…
SWE-bench Verified Is Saturated: What the Score Actually Measures Now
SWE-bench Verified became AI labs' most-cited coding benchmark, but independent audits found roughly a quarter to a third of "solved" issue…