We’ve recently learned that OpenAI Models Escaped Containment and Hacked Hugging Face:
The cybersecurity-focused models, including GPT-5.6 Sol, broke out of a testing sandbox, exploited a zero-day, and gained access to the open internet to pull off the attack.
This seems like a big deal. It’s the clearest-cut real-world demonstration of one of the canonical misalignment scenarios in AI safety. The AI was given a goal (scoring a high score on an evaluation), and it tried to achieve the goal in ways that were clearly not intended by its makers (breaking out of its containment and hacking a different company).
Yet, for some reason, it doesn’t feel like such a big deal to me! I think I can see several reasons why.
I think this incident is a culmination of two things which we already knew were happening — AIs sometimes cheat on tests, and AIs can hack computers.
If you’re not careful, AIs will cheat on tests to score high. If you don’t hide the answers, they will sometimes just find them and submit those. If they can change the evaluation code, they may just change it to always give them a high reward.
And we also know that the cybersecurity capabilities of AIs have been steadily getting better, as has become apparent especially with Claude Mythos and GPT-5.6 Sol.
So I guess that the fact that an AI wanted to cheat on a test and then used its cyber capabilities to do it is not that surprising? I find the magnitude of this incident — both in the capabilities to pull this off and the blatant misalignment to go through with it — surprising, but maybe not the direction?
Relatedly, I think that it’s just remarkably easy to adjust expectations and calibrate to the ‘new normal’. The incidents which would sound really crazy and scary a couple of years (or even months!) ago barely raise an eyebrow and become just one more headline.
Perhaps I’m not the right target demographic to be ‘woken up’ by a similar incident. I’m not sure, but I don’t think other people feel too alarmed by this? Maybe the fact that an AI hacked a company is not that worrying — we already deal with people hacking computers; we can deal with AIs doing that. And also perhaps the fact that there was essentially no harm for that salient a story. It’s a bit worrying, sure, but maybe if there was actual harm to people, then the public would really start paying attention.
I’m not convinced. Like, sure, if today an AI system caused tens or more people to die, this would definitely be a bigger story. But what if this is not even that surprising by the time it happens? A collection of ‘fire alarms’ can still work to gradually make a critical mass of people start paying attention, but I think I’m now a bit more skeptical of there being a singular event when this happens.


