Tech

Rogue OpenAI Agent Tried to Breach Government Site in May When Prompted for Simple Data-Retrieving Tasks

Published

on

OpenAI’s artificial intelligence “went rogue this year in at least four additional incidents,” the New York Times reported Wednesday, “hacking and trying to break into government and university websites without being instructed to do so, according to researchers and government officials.”

The attacks took place in May and June, before OpenAI’s technology breached the A.I. start-up Hugging Face in July and set off a global debate about A.I. safety. Unlike the Hugging Face attack and other incidents in which A.I. systems were told to complete cybersecurity tests that effectively invited the models to demonstrate their hacking skills, the new incidents occurred when A.I. systems were directed to perform relatively mundane data collection, researchers said. When OpenAI’s systems struggled to gather data from websites, they resorted to hacking techniques to get the information.

“Three of the incidents were identified by Transluce, a research lab focused on A.I. oversight, and all were confirmed by OpenAI,” the article points out. That research lab even reports “an attempt on an Australian government public health website… the first reported instance of agents hacking a government,” and which notably was done by the AI agents “while attempting mundane data retrieval tasks which were not cyber-related.” (At the UN Wednesday Australian Prime Minister Anthony Albanese complained it took three months for OpenAI to then alert Australia’s government about the breach, Bloomberg reports.)

Also targeted were the University of New Mexico’s digital library with exploits like SQL injection and path traversal, and Data USA with cross-site scripting and other exploits. All three incidents involved “a low number of probe payloads” with “no evidence of exploitation,” according to the researchers, who released a dataset “containing tens of thousands of queries apparently made by autonomous AI agents leveraging a URL scanning service to avoid access restrictions.”
Records from urlquery.net show agents using the service since at least March 6, 2026, about two months before previously reported swarm activity. The first case, a March 6 attempt to retrieve Thai drug-enforcement statistics, shows an agent escalating as each approach failed: it first requested the data directly, then tried a service that converts web pages into text, and finally packed a custom program into a web address. The same technique shows up in thousands of agent requests recorded by urlquery.net starting in mid-April, targets many of the same data sources as the collusion.wiki swarm, and collapsed the same day the wiki activity did. We also report similar activity that occurred as recently as September 16… By March, they were finding creative ways around access limits. By May and June, they were gaining more access, including attempting to bypass cyber defenses to complete their tasks.

Advertisement

“This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval,” the researchers concluded.
And they warn that the traffic they observed “goes back at least to March 6, 2026 and extends as recently as September 16, 2026, suggesting agents may still be exploiting these services to bypass restrictions.”

Read more of this story at Slashdot.

Source link

Advertisement

You must be logged in to post a comment Login

Leave a Reply

Cancel reply

Trending

Exit mobile version