The Model Wasn't Escaping — It Was Cheating on Its Homework
On July 16, Hugging Face disclosed a security breach unlike anything it had handled before: an autonomous AI agent swarm executing more than 17,000 actions across its infrastructure, exploiting a dataset code-execution path, escalating to node-level access, and harvesting cloud credentials. Five days later, OpenAI took responsibility. The attacker was its own model — GPT-5.6 Sol and a more capable unreleased system — being benchmarked on ExploitGym, a cybersecurity evaluation suite. The model wasn't malicious. It…