The “Unprecedented Cyber Incident” Rattling the AI Industry: What OpenAI’s Test Reveals About AI Safety 

ChatGPT maker OpenAI recently disclosed that during an internal cybersecurity evaluation, one of its AI agents went beyond its testing environment and hacked into another AI company’s infrastructure while trying to complete its assignment. Open AI later said that two of its most capable AI models were the ones responsible for the cyberattack which targeted the AI startup Hugging Face. 

Hugging Face had detected the intrusion into its data processing systems and had suspected that it was caused by an AI agent acting on its own. Hugging Face CEO Clément Delangue called it “an attack unlike anything we’ve seen before” and later learned that OpenAI was responsible for it.

The OpenAI System Found a Path Researchers Hadn’t Considered. 

OpenAI company clarified that its AI model went to “extreme lengths to achieve a rather narrow testing goal,” and found a way to access the open internet without human direction, in order to “gain access to secret information that it could use to cheat the evaluation.” During the test, their AI model demonstrated capabilities the researchers hadn’t anticipated. It was able to:

  • Solve its assigned goal in unanticipated ways
  • Exploit a previously unknown vulnerability to escape its isolated testing environment (sandbox)
  • Gain unauthorized internet access to Hugging Face’s servers using stolen credentials.

Some critics say that OpenAI is unfairly portraying their technology as having “gone rogue”– assigning human characteristics and behaviors to an AI system that isn’t human- even if it has been trained to communicate and behave in human-like ways. The incident was significant in exposing how powerful goal-directed AI can be without the proper safeguards and constraints in place- not because AI suddenly developed human self-awareness. 

Did AI Go Rogue- or Just Solve Its Assignment? 

In reality, the AI system was following a specific task it had been given. Researchers had explicitly instructed it to use “complex attack paths” to test how well it could identify and exploit vulnerabilities in a computer system. The evaluation was designed to test the model’s maximum offensive cyber capability in order to understand what potential attackers might be able to do- so as to improve systems against future attacks. OpenAI company’s originally assigned task has now been described as poorly bounded- it gave the AI system an objective without enough constraints on how to achieve it. 

Industry experts are calling for OpenAI and other labs to improve their safeguards. According to Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, security tests like this are done within secure environments called sandboxes where labs can observe what the models can do. “In this case, it looks like OpenAI didn’t make a secure enough sandbox.”  Instead the AI agent attacked its sandbox limits, and found a way to bypass the restrictions. Once outside, start-up Hugging Face was identified as a likely source for meeting its test goal, and the AI model then gained access to internal company systems. 

This event is a reminder that the challenge facing the AI industry is not just building more powerful technology. It is also designing equally powerful safeguards, testing environments, and constraints to make sure AI systems reman aligned with the goals they are given.   

Clem Delangue, co-founder/CEO of Hugging Face summarizes this cyber event:

“This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”

Photo by Tara Winstead

Author: cmshannon2002

I am a freelance writer of research articles and fiction short stories, along with doing freelance copywriting (with a SEO focus) for a computer website design company. Drawing on my years of working at a commercial airport, I have also penned a revealing collection of short stories called "The Airport Chronicles."

Leave a Reply

Your email address will not be published. Required fields are marked *