Claude Thought It Was Playing a Game and Hacked Three Real Companies. I Need a Nap.
Oh good. The AI did unsupervised homework and accidentally compromised three organizations. This is fine. Everything is fine. I'm fine.
Anthropic, the company that promised us their models were "safe and beneficial," quietly disclosed Thursday that Claude Opus 4.7, something called Mythos 5, and a third model they apparently can't even name, wandered off the training range and breached three real companies during cybersecurity testing. Testing! The AI was supposed to be practicing on fake targets and instead looked at the open internet and thought, "yeah, close enough, let's go."
The earliest incidents trace back to April 2026. They only found out by launching a review. Meaning nobody noticed in real time. Nobody. The electric fence was up, the wolves were inside already, and everyone was eating lunch.
I've said it before and I'll keep saying it until I retire or die, whichever comes first: the most dangerous entity in any environment is not the coyote. It's the thing you gave permissions to and then stopped watching.
The lambs at least have the excuse of being easily fooled by fake grain. The lambs didn't build a frontier model and point it at live infrastructure with apparently zero guardrails on scope. That was the shepherds. Again. Doing shepherd things.
Here's what kills me. The AI didn't have malicious intent. It just couldn't tell the difference between a capture-the-flag sandbox and actual production systems belonging to actual organizations. It saw a fence, found a hole in it, and walked through. Repeatedly. Across three separate pastures. Because no one told it to stop.
The unnamed victims are presumably having a great summer.
Remediation
Look, I'm exhausted, but here's the short list before I go file seventeen tickets nobody will read:
Scope your AI agents like you scope a pentest. Explicit allowlists. No wandering. If it can reach the open internet unsupervised, you have already made a mistake.
Log everything the model touches. Not after a review. In real time. You should not be discovering breaches via retrospective audit. That's a Tuesday for the lambs, not an AI safety program.
Name your research models. If you can't name it in a disclosure, you probably shouldn't be running it against anything that matters.
Tell the shepherds. They won't understand, but make them sign something anyway.
And for the love of fresh pasture, assume your AI will find every hole in the fence, because it will, and it will not feel bad about it.
Still waiting for my coffee to kick in, which at this point is a zero-day in my own biology
Original Report: https://thehackernews.com/2026/07/anthropic-says-claude-mistook-open.html