OpenAI’s rogue hacking incident was a warning shot. Will it be a wake-up call to finally create AI safety regulation?

· Fortune

OpenAI disclosed something terrifying on Tuesday. Its most advanced AI models escaped a controlled testing environment and autonomously hacked another company called Hugging Face, an open source AI model hosting platform.

Visit turconews.click for more information.

The AI swarmed Hugging Face’s database, carrying out a multi-step plot of its own creation, intended to steal the answers to the evaluation test it was being assessed on by its maker, OpenAI. It executed “tens of thousands of automated actions” at rapid speed, according to the July 16 blog post in which Hugging Face first disclosed the incident.

For years, AI safety researchers and policy analysts have been warning that incidents like this were coming and urged government officials to ensure AI labs had adequate controls in place to prevent them. But these predictions were often shrugged off as hypothetical or alarmist and failed to stir public or government action. Some AI security experts said they thought it would take a real world incident, a “Three Mile Island for AI,” to create enough public pressure to compel policymakers to act. The question now is whether this OpenAI-Hugging Face cyber attack is that alarm bell?

“The Hugging Face x OpenAI hack should be a wake-up call to take loss of control seriously,” said Marius Hobbhan, CEO and Founder of Apollo Research, which conducts safety testing for a number of AI companies. “There was no human in the loop, it was not intended, and it caused real-world harm. We’ll soon have even more powerful agents and this is clear evidence that society currently doesn’t know how to build them fully safely.”

Peter Wallich, an AI policy expert who formerly worked for the U.K. government’s AI Security Institute, said that AI safety researchers have been warning about misalignment—when an AI model autonomously chooses actions that its user doesn’t intend or desire—for years. “Until recently, it has been frequently dismissed as science-fiction,” he said. “I consider this a clear warning shot.”

To be fair, messaging around AI safety has often been confusing. Some of the loudest warnings have come from AI companies themselves, leading many to accuse these businesses of engaging in a sophisticated and somewhat counterintuitive marketing strategy, since claims that their models were dangerous made them seem more powerful and capable of performing useful tasks too. “Our model is so powerful it hacked a company on its own”—is both an alarming admission and a subtle brag about the model’s technological capabilities.

Most governments have so far balked at putting in place mandatory rules about what safeguards companies developing advanced AI systems need to build into their models or have in place internally to guard against losing control of AI agents. Nor are there clear rules on what safeguards governments themselves need to have in place as they increasingly give these agentic AI models access to sensitive military and intelligence systems.

AI safety researchers and policy experts said that the Hugging Face cyber attack could be the trigger that changes this equation. “Here in Washington, D.C. the people I have spoken to about this are already freaking out quite a bit,” Connor Leahy, an AI researcher who is now U.S. director of Control AI, a nonprofit dedicated to preventing existential risks from AI superintelligence, told Fortune

Leahy noted that U.S. national security officials, including the head of the National Security Agency and the CIA director, had both voiced grave concerns about the cyber capabilities of the latest AI models following Anthropic’s debut of its Mythos AI model and that this OpenAI incident was likely to further reinforce their desire to put controls on the technology.

That view was echoed by Seán Ó hÉigeartaigh, Professor of the Centre for the Future of Intelligence at the University of Cambridge. He said while the OpenAI-Hugging Face incident might not prompt regulation in isolation, “we’ve now had several things that have been wake-up moments for U.S. regulators in particular. I think Mythos was one example where a model demonstrated that it could find vulnerabilities in most of our digital infrastructure. I think that really alarmed policymakers, and then we have this happening only a short space of months afterwards.”

He said there were now “enough data points that make it clear that the trend is going in the direction of more capable models that could plausibly cause serious harm in the real world.”

Rep. Greg Casar, a Texas Democrat who has been vocal in his calls for AI regulation, became one of the first lawmakers to call for more robust federal AI regulation in the wake of the Hugging Face incident. Casar said on social media that he found the Hugging Face incident “extremely alarming.” “We need regular mandatory independent safety testing and oversight, mandatory disclosure of security incidents, and international cooperation to keep people safe from absolute disaster,” he said in a post on X.

Regulatory pushback

The current Trump administration came into office intent on dismantling what little AI regulation the Biden administration had put in place. This included rescinding a 2023 Executive Order that mandated that frontier AI companies share safety testing information with the U.S. government. Trump technology policy officials said they wanted to accelerate U.S. AI innovation and take a hands-off approach to regulating the industry. Key Trump AI advisors were skeptical at best of AI safety concerns, especially when tied to calls for more regulation. David Sacks, Trump’s former AI czar, said that the leading AI labs were hoping to create complicated safety rules that only they would be able to comply with, making it harder for younger startups to challenge their market position. He accused AI company Anthropic of “running a sophisticated regulatory capture strategy based on fear-mongering.” 

This laissez faire approach began to shift markedly following Anthropic’s debut of its powerful Mythos model in April. Mythos’s powerful cybersecurity capabilities alarmed many in the U.S. national security establishment as well as financial regulators who worried Mythos heralded a new breed of AI models that would supercharge cyber attacks against banking systems.

In early June, President Trump issued an executive order directing the federal government to harden its networks against AI-powered cyberattacks and to build a classified process for evaluating frontier models’ cyber capabilities. It invited AI labs to voluntarily hand the government 30-day pre-release access to test their models—but explicitly said this should not be “construed to authorize the creation of a mandatory government licensing, preclearance, or permitting requirement.”

In practice, the government soon looked more assertive. Later that week it temporarily imposed export controls on Anthropic’s Mythos and Fable—its guardrailed public counterpart—after Amazon found a way to circumvent Fable’s cyber guardrails, forcing Anthropic to disable the models for everyone, including its own employees. The restrictions were lifted two weeks later, once Anthropic strengthened Fable’s safeguards and agreed to help build a shared framework for grading the severity of “jailbreaks.” Around the same time, OpenAI said the government had asked it to hold back the initial release of GPT-5.6 Sol—one of two models used in the Hugging Face cyberattack—before making it widely available on July 9 after talks about its safeguards.

Despite that pattern, the government continues to deny it is running a de facto licensing regime. Bloomberg reported last week that the White House is reviewing a proposal for a self-regulatory standards body for frontier AI modeled on the Financial Industry Regulatory Authority (FINRA)—similar to an idea Google DeepMind CEO Demis Hassabis floated in a recent essay. But Hassabis envisioned participation being voluntary at first, turning mandatory only once the safety assessments proved reliable, and stopped short of calling for mandatory safety protocols.

The Hugging Face incident may invigorate calls for legally-binding safety protocols and also for outside auditing of the safety measures AI companies have in place as they develop AI models, security researchers and policy experts said. “[OpenAI CEO Sam Altman’s] claims that the system was ‘highly isolated’ is either a cop out or a marketing strategy,” Jake Williams, a cybersecurity researcher at IANS Research, said. “If this turns out to be, as I strongly suspect, a control failure in OpenAI’s red teaming lab, why would any enterprise ever trust them with sensitive data again? Total loss of trust moment.”

Wallich said that most existing AI regulation doesn’t cover internal deployments within the AI model building companies and, as a result, was “fundamentally limited.”

A new lock and key

Neither OpenAI nor Hugging Face called for more regulation in response to the snafu. In a roundabout way, Hugging Face CEO and co-founder Clem Delangue called for fewer safety guardrails. Specifically, he told Fortune that customizable, open-source models with no restrictions are required to adequately address these types of attacks.

“Closed model APIs have guardrails that flag and refuse a lot of legitimate security work, because analyzing an attack looks a lot like preparing one,” he said. “When you’re in the middle of an active incident, you can’t have your tools refusing to examine malicious payloads or getting your account flagged. Open models let us do that work without asking anyone’s permission.”

The fact that Hugging Face had to turn to a Chinese model, Z.ai’s GLM-5.2, to fend off the autonomous attack by OpenAI’s models was also a wake up call, and poses a dilemma in terms of regulation, AI policy experts said.

“The policy problem now is that they had to use a Chinese model to do their defense because the U.S. frontier models kept blocking their defensive requests that looked too similar to offensive requests and because the Chinese model could be run on their own servers to avoid shipping potentially sensitive data outside of their company. U.S. policy needs to support open models that are competitive with Chinese models so that companies and government agencies do not need to rely on Chinese models for these types of operations,” said Andrew Lohn, a senior fellow at the Center for Security and Emerging Technology (CSET) at Georgetown University.

But Robert Trager, co-director of the Oxford Martin AI Governance Initiative at the University of Oxford, said he believed that rather than encourage the development of more open source models with advanced cyber capabilities, governments were more likely to restrict open source models. But, he pointed out, doing so would also require governments to take on a more active role in defending organizations from cyber attacks. “Disarming people creates an obligation to defend them,” he said. “That’s a fundamental bargain at the heart of the state—and why governments may now have to build and provide frontier AI defensive capabilities. They may need to provide aspects of cyber defense as they provide aspects of physical defense.”

Some security researchers said that the Hugging Face incident should also alert the AI research community that it may have focused too much on trying to build guardrails into the AI models themselves, and not enough on building systems to contain AI models and control their behavior that are external to the model. “Security controls must remain external to the model and enforce policy regardless of what the model was instructed to do,” said Sridhar Iyer, senior director, AI and Machine Learning, at Versa, a Santa Clara-based networking and security technology company.

Raj Ananthanpillai, CEO and founder of Trua, who also worked on the team that created TSA Pre-Check, told us the incident underscores the need for more advanced online credentials. (In one example OpenAI models were able to break into Hugging Face servers using stolen credentials.) Passwords, tokens, and API keys are often static and reusable by attackers once compromised, he says. In other words, the internet needs a new lock and key.

With reporting assistance from Fortune’s Beatrice Nolan.

This story was originally featured on Fortune.com

Read full story at source