Kate Conger
San Francisco: OpenAI said on Tuesday (California time) that two of its artificial intelligence models misbehaved and were successfully hacked into Hugging Face, a digital library of AI technology popular among developers.
The incident, which happened last week when OpenAI was testing the cybersecurity capabilities of its systems, showed the kind of science fiction possibilities that AI companies warned would soon come true.
AI labs like OpenAI and Anthropic have in the past year released AI models that have been optimized to uncover cybersecurity problems, warning that their technology could introduce new risks by finding holes in corporate computer networks faster than defenders could fix them.
OpenAI’s disclosure on Tuesday is a sign that such security incidents are starting to happen, and even savvy AI companies may not be ready to deal with them. New AI systems can take many steps, find ways around obstacles and find new ways to attack the network, said Alex Levinson, a cybersecurity consultant who focuses on self-sufficiency.
“That’s a real threshold, and it’s going to be a normal part of the security environment,” he said.
The Hugging Face attack began when OpenAI tested a combination of two of its models, GPT-5.6 Sol and a more powerful, yet-to-be-released model, to see how well those models could integrate cyber vulnerabilities into successful cyber attacks, OpenAI said in a blog post about the incident.
The test was designed to put designs in a safe testing environment, known as a sandbox, OpenAI said. But those models found vulnerabilities that allowed them to escape the sandbox and connect to the Internet. They then focused on Hugging Face because they hypothesized that the library, which contains millions of AI models, might hold clues about how to pass the test.
“It seems to me that OpenAI didn’t create an adequate sandbox as a testing environment,” said Dierdre Mulligan, a professor at the School of Information at the University of California, Berkeley who focuses on security and AI systems. He questioned whether passing the test was worth the potential damage of an AI-style escape into the internet.
“What do we get, and if this is the only way these experiments can be set up, what are the risks?” He said.
OpenAI said it was working with Hugging Face to fix the issues that led to the attack.
“We are treating this as an unprecedented, high-risk cyber incident, and we are responding accordingly,” OpenAI said in its blog post. “We implement strict controls on infrastructure configuration at the expense of research speed when vulnerabilities are identified.”
Hugging Face said last week that it had discovered the attack and knew it was caused by an autonomous system, but did not say at the time that OpenAI was involved.
Clem Delanggue, CEO of Hugging Face, said Tuesday that his company worked closely with OpenAI over the past 24 hours to address the attack.
Delanggue said in a statement that he was “grateful for the cooperation” with OpenAI following the hack. “This event, perhaps the first of its kind, confirms what we have long believed: AI security will not be solved by any company operating in secret,” he said.
AI models have proven to be adept at programming, and that has made them valuable to hackers and people in charge of protecting computer networks.
In April, Anthropic released a cybersecurity-focused model called Mythos and made it available to just a small group of organizations to protect themselves against cyberattacks. OpenAI recently introduced its cybersecurity model and made it available to a small group of organizations to develop their own defenses, before rolling it out more widely. And on Tuesday, Google said it has also developed a model that focuses on cybersecurity and is releasing it to a small group of test partners.
(The New York Times has sued OpenAI and Microsoft, alleging copyright infringement of content related to AI systems. The two companies have denied the claims.)
Richard Barnes, an independent security researcher who has worked with Mythos, said the cybersecurity industry faced a similar challenge about a decade ago, when new tools called fuzzers made it easier for attackers to break into online systems. Technology companies began using tools to check their own systems for vulnerabilities and were eventually able to prevent many attacks.
Companies must now take a similar approach to prepare for AI attacks, Barnes said, “before those vulnerabilities are found and exploited by bad guys who have access to these tools.”
New York Times
Start the day with a summary of the day’s most important and interesting stories, analysis and insights. Sign up for our Morning Edition newsletter.




