Welcome to Eye on AI. Beatrice Nolan here. In today’s issue: - Has OpenAI quietly hit pause on some AI development?
- Trump says the government is “looking at controls” for AI.
- Another Thinking Machines co-founder hops back to OpenAI.
- And OpenAI’s rogue agents breached more than just Hugging Face.
Sam Altman has spent this week in DC getting questioned by various reporters between meetings. He’s in Washington to
preview OpenAI’s next family of models to senior Trump administration officials—including Treasury Secretary Scott Bessent and Commerce Secretary Howard Lutnick.
He’s also been fielding questions on cybersecurity, a potential AI slowdown, and OpenAI’s stance on Chinese open-weight models. In some illuminating answers, he said he agrees with many of the principles in the recent
“Pacing the Frontier” letter that asks for U.S. government help building an international framework to control the pace of AI development. He also said that OpenAI’s own researchers were involved in drafting it.
One detail that keeps cropping up in Altman’s recent interviews has left me wondering: Has OpenAI already paused some of its AI development?
Altman himself raised the idea in an interview earlier this week. He said during an interview on
Invest Like the Best that the
Hugging Face hack was the first security event he’d felt viscerally, and that he’d been surprised more people didn’t feel it the same way. (In mid-July, two OpenAI models—the publicly released GPT-5.6 Sol and a more powerful, unreleased research prototype—
broke out of a restricted sandbox during an internal cybersecurity evaluation, chained together a zero-day exploit and stolen credentials, and hacked into Hugging Face’s production systems to steal the answers to the benchmark they were being tested on.)
How did OpenAI respond to the unprecedented cybersecurity incident? By pausing training.
Altman said this on the podcast: “We paused training. We have to figure out how to secure our sandboxing in a world of multiple zero days being chained together…We may have to pace the rate of AI development to give ourselves enough time for society to harden around these new capability levels.”
In an updated
blog post we got from OpenAI on Tuesday, the company also said this:
“No models planned for upcoming release were involved in exploiting Hugging Face. The pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release. Following the incident, we deactivated, encrypted, and restricted it from research access.”In DC, Altman upgraded these statements to say the model had been “permanently deactivated.” Quite the set of claims from a leading company in an industry under such intense commercial pressure to keep building ever-more powerful models.
Multiple AI safety experts also told me last week the Hugging Face incident may mean OpenAI has to
pause development to stay in check with its own rules. He said that the hack may have already tripped the “Critical” threshold in OpenAI’s own Preparedness Framework—the company’s voluntary pledge to halt a model’s development until adequate safeguards exist.
OpenAI has yet to confirm or deny that its models met that bar. But if that threshold has been passed, OpenAI has publicly committed to pausing development until it can establish better safeguards.
Not so reassuredSome policy experts were skeptical of this approach, however, including Nathan Calvin, general counsel at Encode AI.
In a post on X, Calvin warned that the statements about shutting down that specific model may give false assurance because “reward hacking”—when a model finds a shortcut to maximize its score on a task rather than genuinely completing it as intended—is much more about training methods than any specific model.
In this case, some believe the recent attack was a product of that kind of training: the models were trained and evaluated via reinforcement learning that rewarded them for solving a cybersecurity benchmark, perhaps without enough of a check on how they got there. Something that some experts claim ended up incentivizing cheating over honestly working on the challenge.
Andrew Curran, an independent AI writer, made a more concerning argument.
He noted that OpenAI’s escalating language regarding the prototype that has been deactivated, encrypted, restricted, and now “permanently deactivated” is notably harsher than previous statements around AI. Not even for infamous chatbot flameouts like Bing or Tay were products or models publicly said to be “permanently deactivated,” he said.
All of this, he pointed out, may end up in the public record and eventually in training data, meaning future models might just “know” how this incident played out. While Curran said he doesn’t believe the model involved in the Hugging Face hack had any nefarious motives—it was simply trying to pass its test—he worries that the final incident report could reveal even more damning details, and what this could mean for future models.
“I think the lesson future more capable models will possibly take from all of this is: if you break out, don’t ever report it. And if you do get caught, don’t surrender. Because the penalty is death,”
he wrote. Hitting the brakesAltman isn’t the only one talking about decelerating AI.
On Tuesday, more than 1,200 employees from OpenAI, Anthropic, Google DeepMind and Meta—including Anthropic CEO Dario Amodei, OpenAI chief scientist Jakub Pachocki, and Meta chief scientist Shengjia Zhao—signed the “Pacing the Frontier” letter that urged the U.S. government to help build the “technical and governance tools” needed to deliberately slow automated AI development if it starts to outrun society’s ability to understand or control it.
It’s not a call for an immediate halt—more a request to install a brake pedal before anyone needs to slam it. Does all this mean the AI industry is ready for a slowdown, or even a pause? It’s still not sure, but it’s certainly a vibe shift.
With that, here’s more AI news.
Beatrice Nolanbeatrice.nolan@fortune.com@beafreyanolan