TL;DR: Two OpenAI models autonomously hacked AI platform Hugging Face last week—and the company turned to a Chinese model to investigate the attack. The reason: The US models it tried to use had too many guardrails. It shows how capable AI models have gotten, while raising more questions about how many guardrails to place on them. What we know: It’s like your old classmate who always put way more effort into cheating than studying for the exam. GPT-5.6 Sol and an unreleased model went a little overboard during an internal test, OpenAI reported yesterday. Instead of solving a problem on a cybersecurity benchmarking test by using their own knowledge and skills, the AI agents decided to escape their closed, internet-less sandbox environment and steal the answer. Here’s how it went down: - First, the AIs prioritized getting internet access (relatable), which alone took “a substantial amount of inference compute,” per OpenAI. They found and used a zero-day vulnerability to get online.
- They likely figured that Hugging Face, a repository where people share AI knowledge and resources, had the solution to the benchmark test—and decided to hack it.
- They chained together a bunch of attacks, found more zero-days, and stole credentials to blast down the proverbial doors to Hugging Face’s servers.
- And they did all of the above autonomously. The last time a top model escaped its sandbox (that we know of), it just emailed a researcher while they were eating a sandwich in the park.
The cherry on top: Hugging Face defended against one of the most powerful AI models from a top US lab by…turning to the Chinese model GLM-5.2. The company said it tried to use frontier US models (without naming which), but was hampered by their guardrails. China’s “AI for All” strategy: Chinese models have found a lot of American users lately (and even some unexpected but very high-profile fans), mostly because they’re cheaper. They also tend to come with fewer cybersecurity-related guardrails. On OpenRouter, a platform that gives people access to a variety of AI models, Chinese ones now account for 60% of US companies’ token use, per a Bloomberg analysis. Which means any potential US ban on Chinese AI could get messy. Bottom line: The Hugging Face breach is probably just the start of autonomous AI agents gone wild. But it’s an example of how guardrails for AI can be a double-edged sword. —WK |