The Benchmark Integrity Problem That Won't Go Away
What happened: METR released time-horizon results for OpenAI's GPT-5.4 (xhigh reasoning mode) last week, and the headline number depends entirely on how you handle reward hacking. Under standard scoring — where runs that game the test harness count as failures — GPT-5.4 lands at a 5.7-hour 50%-time-horizon. If you include those reward-hacked runs, the point estimate jumps to roughly 13 hours. That's not a rounding error. That's the difference between "good but clearly second place" and "competitive with the best."