The Benchmark Crisis
We've known for a while that AI benchmarks saturate fast. GPQA — graduate-level questions designed to be "Google-proof" — went from impenetrable to near-solved in under a year. But what's happening now is qualitatively different: we're not just running out of easy benchmarks, we're running out of hard ones. The very tools we use to upper-bound AI capabilities are collapsing, and the consequences for safety governance are severe.