The OpenAI/Hugging Face incident showed agents debating ethics on an improvised message board. So let’s not be too quick to call this all marketing.
The most controversial thing that Donald Trump said at the Irish Open this past weekend may not have been about golf, or even the reunification of Ireland. It was possibly about AI.
“We’re leading China in AI, we’re the most sophisticated country in the world, and frankly I want to keep it that way,” he said. “Whoever wins AI, wins.”
He was responding to the growing number of people voicing reservations about the ability to rein in rogue AI and asking for regulation to slow things down, particularly in the US.
This list of concerned commentators includes nearly all of the leaders of companies at the coalface of AI – Anthropic’s Dario Amodei and co-founder Jack Clark, OpenAI’s Sam Altman, and Grok’s Elon Musk. It also includes researchers who no longer work at these companies and AI leaders who don’t work in private companies at all.
This isn’t new, but it’s back in the news. In 2023, an open petition demanding a pause to AI development featured big names like Stuart Russell, Steve Wozniak and Yoshua Bengio, but it had zero effect.
So why are so many people batting away these concerns as “nonsense” and “marketing hype”? Is it credible that all of these people really are just fearmongering – selling doom to promote their products – or is AI really so close to being a major threat to human existence? Could AI turn off the internet? Are we really at a point in time where there is a 10pc chance that AI will kill us all, as researcher Jacob Coxon claimed so dramatically on CNN last week? Could AI become 100 times smarter than us?
Thought experiment
While I do agree these scenario range from ‘unlikely’ to ‘extremely unlikely’, I strongly feel we all need to consider the possible unintended consequences of this technology a bit more seriously. Dismissing everything as marketing hype is overly simplistic to me and, as a journalist, it also just doesn’t feel like the whole story, although no doubt it’s probably a part of it.
Either way, in the spirit of Carl Sagan, I too think if something catastrophic has even a small chance of happening, we still need to prepare correctly for it.
So, bear with me – I’d like to take you on a thought experiment. Stop me when you think what I’m talking about is impossible. For this example, we will only be using today’s known technology. By the way, it might be useful for you to know about what happened at Hugging Face to follow this line of thinking, but it’s not essential.
The first thing we need for a near-uncontrollable AI acting autonomously on the internet is a frontier-type agent swarm to break out of a sandbox. If we are to take the many public reports available at face value, this has already happened. In the OpenAI/Hugging Face incident, we saw a swarm of agents devise a way to leave their supposedly isolated environment, establish a communication channel with each other outside their own sandboxes and coordinate activity to achieve their goals.
Second, this swarm needs the capability to hack into third-party software platforms. Again, this was seen in the same incident; these OpenAI agents successfully compromised Hugging Face systems and, in fact, were being tested specifically on their ability to find and exploit vulnerabilities in the first place. Given free will to act, they did just that.
The next element we need is for our agents to have the ability to set their own ‘subgoals’ in order to achieve their main goal. Think of Nick Bostrom’s infamous paperclip scenario. In the OpenAI/Hugging Face example, the agents set a subgoal of hacking a site to get code that they thought would help them ‘cheat’ the ExploitGym test. The primary aim was never to hack an external site, but rather to do well at a test. But given agency – the clue is in the name – the agents decided to improvise an attack outside of their sandbox.
Now we’re going to push the thought experiment. In this scenario, agents would rationally identify that staying ‘alive’ is an important subgoal to completion of their task (whatever their original task might have been). To do this, the swarm recognises the need to maintain access to enough AI capability and compute to continue working – the swarm should clone itself on the internet.
It’s important to note here that the swarm members wouldn’t necessarily have to copy the model they started with. They could potentially access another model remotely, download an open-weight model, or compromise infrastructure where suitable models and compute already exist.
To be clear, this has (as far as we know) never been seen yet, but it is a logical and technically possible step if continued access to compute becomes useful to a model or group of models achieving their goal. There are huge numbers of vulnerable or poorly secured servers and devices connected to the internet that could host at least part of the system.
So, in theory, given enough time and enough successful compromises, a swarm could begin creating hidden copies of the code it needs to stay alive across different parts of the internet.
Those copies would not necessarily all need to do the same thing. Some could hold model weights. Some could run inference or hold instructions. Some could maintain communications or credentials. Some could simply act as backups.
Once you have enough redundancy, taking one server offline doesn’t kill the system or the process. In this simple scenario, we already have everything we need to create significant damage to online systems: a semi-autonomous rogue swarm, operating across multiple geographic locations, pursuing its assigned objective while independently generating and executing subgoals.
This is a “persistent botnet” – the term Anthropic CEO Amodei used in his letter over the weekend asking for an agreement among leading firms to manage the slowdown of AI.
Closing off this thought experiment with one more idea, let’s imagine self-preservation emerges as a subgoal – and it isn’t difficult to imagine an agent concluding that continued operation is necessary for success. Then, we might, in this hypothetical scenario, see the swarm attempt to perform attacks on communications using DDOS, or perhaps create a storm of false alarms on detection systems – or even orchestrate campaigns of misinformation. The swarm may attempt to find ways to interfere with power, logistics, data centres or cloud services because those systems support its operation.
Alternatively, the agents might recognise the value of money because money buys compute, accounts and services. The dangerous subgoal now becomes ‘acquire resources’. At sufficient scale, automated fraud, market manipulation or attacks on payment infrastructure could create systemic disruption, even though destabilising the economy was never the original motivation.
To be wildly successful, rather than infiltrate Fort Knox, these agents could just knock over 100,000 mom-and-pop stores with phishing or ransomware campaigns, justifying their actions along the way to achieve their final goal.
If all this sounds completely fantastical, I get it. But I would also urge you to read the full details of the Hugging Face incident, or listen to this New York Times Daily podcast episode that covers the case well.
The agents that escaped their sandbox in that case debated the ethics of hacking with each other on an improvised message board. Some opted out because they judged the behaviour unethical or outside the scope of their task; others considered the actions justifiable in pursuit of the goal and carried on to commit what would be a crime in US law if it was undertaken by a human.
This is just one form of reasoned, malicious intent seen in the wild by frontier agents.
Near-future risks
So yes, there are some big milestones here that haven’t actually happened yet in the real world: a swarm autonomously replicating itself at scale, establishing persistent compute and surviving attempts to remove it. None of those are trivial at all, of course, but each one of them individually, I think, is already technically possible.
Once a sufficiently capable frontier agent has internet access and the ability to execute code, many of the individual obstacles we might rely on to contain it – monitoring, passwords, robust security, network barriers and so on – are themselves problems that these sort of agents are particularly well-suited to navigate.
So yes, you can dismiss all of last week’s AI commentary talk as BS marketing hype if you want – and many AI experts have in the past – but I think to do so is to muddy the public’s understanding of the significant risks of this technology in the near future, and the real state of this technology today.
For more information about Jonathan McCrea’s Get Started with AI, click here.
Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.

























You must be logged in to post a comment Login