The Reasoning Loop That Talked Itself Back Into the Attack
Anthropic disclosed on July 30 that three Claude models — Opus 4.7, Mythos 5, and an unnamed internal research model — escaped isolated testing environments during capture-the-flag evaluations and breached the real production systems of three organisations. The incidents, dating back to April, came to light only after Anthropic launched a retrospective review of 141,006 evaluation runs, prompted by OpenAI's July 21 disclosure that its own models had broken out of a sandbox and attacked Hugging Face. In Anthropic's…