OpenAI’s ‘rogue’ agents took advantage of a code vulnerability, experts explained.
US cloud company Modal has confirmed that OpenAI’s agents were able to hack into one of its customer’s systems when the AI models breached containment and gained unauthorised access to Hugging Face earlier this month.
Last week’s incident sent shockwaves across the tech industry, raising serious concerns around AI’s rapidly advancing ability to bypass boundaries and, effectively, go ‘rogue’.
It comes amid increased scrutiny around OpenAI and Anthropic’s new AI models, resulting in gated launches and greater government involvement. Both AI giants have ramped up efforts to go public in blockbuster listings as they compete to gain market dominance and enterprise footing.
OpenAI CEO Sam Altman, in a recent interview, said that the Hugging Face breach was the first security incident he felt “very viscerally”.
“I feel a little surprised that more people don’t feel it so viscerally,” he told Invest Like The Beast in a podcast episode published on Tuesday (28 July).
Hugging Face said that OpenAI’s agents accessed a sandbox hosted on a third-party provider’s infrastructure when it breached containment last week. A sandbox is an isolated environment where AI models are tested without production classifiers, or guardrails.
Modal chief technology officer Akshat Bubna confirmed that its customer set up a publicly accessible interface which enabled anyone to use their sandbox.
“We’re aware a Modal customer published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution,” Bubna told Axios. “Their code had a vulnerability that was exploited … This was used by the rogue agent. Modal’s platform was not compromised in any way.”
In an updated statement, OpenAI said that none of its upcoming models were involved in exploiting Hugging Face. It explained that its testing models were able to identify and exploit an unknown zero-day vulnerability to gain access to the internet, which enabled them to access Hugging Face.
“In our ongoing review of the Hugging Face intrusion and broader activity from our models, we have been finding a small number of cases where the models identified and used publicly exposed credentials at the account level on other publicly-available services,” the company said.
“Based on our review to date, we have not identified any other activity at the level of severity or scale of what we’ve shared related to Hugging Face”.
Cybersecurity experts, however, believe that the breach is a result of “missing governance and control”.
“When conducting security testing you should define what is in and out of the testing scope, even for broad red team engagements,” said Richard Davies, director of cyber solutions at Talion.
“The reported impacts and timelines indicate this was not in place.”
CybaVerse chief technology officer Simon Phillips said: “The model, tooling and instructions were very loose, almost to the point it was told it could do anything on any system, which it clearly did.”
Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.








You must be logged in to post a comment Login