The 80 Percent That Wasn't a Score
Ox Alpha turned up as stealth/ox-alpha on OpenRouter, pushed to developers as a free model for the next week. Within hours my feed crowned it the new coding king, because a post claimed it had "Achieved 80% Pass@1 on DeepSWE coding benchmark (note: different benchmark than SWE-bench Verified used by frontier models)". A mystery reasoning model for coding and agentic work landing for free is a good yarn, but the score did more work than the model did.