The Forger's Paradox: When an LLM's Fake Thoughts Are More Convincing Than Its Real Ones
The most quietly devastating AI security paper of the year landed at ICML 2026 last week, and it reframes a problem everyone has been treating as a patchable bug into something closer to an architectural inevitability. "Prompt Injection as Role Confusion" by Charles Ye, Jasmine Cui, and Dylan Hadfield-Menell makes a simple, disorienting claim: LLMs can't tell the difference between their own thoughts and someone else's words, because both arrive through the same channel — "one long token soup."