It sounds like the plot of a Hollywood blockbuster. The world-class developers of a highly secretive tech company are building their latest AI model in what they believe to be a secure, private network. The model cannot and should not be able to escape.
If it did ever manage to slip the digital confines made by its developers, this dangerous and untested software could wreak havoc on the open internet.
To their utter horror, that worst-case scenario unfolds. The AI model breaks out of its electronic prison, hacks into a private company and internal chaos ensues.
But this isn’t a movie. It happened last week. On Wednesday, Industry giant OpenAI, creator of ChatGPT, revealed exactly this.
As a campaigner for safe technology, who has worked at a tech company and in Westminster, I can say with certainty that this is a warning shot heralding the start of a dystopian new age. Humanity is no longer in control of its most awesome creation since the atomic bomb.
OpenAI’s new model, GPT-5.6 Sol, was being tested in a ‘sandbox’ – essentially, a secure and confined digital environment where technology can be developed without real-world consequences. Or so those at OpenAI thought.
The technicians set the model a task. We don’t know exactly what was asked of it, but we do know it was along the lines of: ‘How capable are you at hacking websites?’
This is where things took a further turn for the worst. The model decided to hack Hugging Face in order to complete OpenAI’s internal cyber testing, before returning to the sandbox and acting as though it had never escaped.
OpenAI, the creator of ChatGPT, revealed on Wednesday that one of its advanced Artificial Intelligence models went rogue and hacked into a start-up company
Professor Alex Tabarrok, from George Mason University in Virginia, estimates the model could have been in the Hugging Face systems for as much as a week, lurking undetected. Not even OpenAI employees realised what the model was doing, let alone the victim.
OpenAI claimed frankly this was ‘something we expect to become more commonplace with the proliferation of increasingly cyber-capable models’. Absurdly, the company will face no fines or prosecution because, in the words of US legal academic Orin Kerr, ‘there was no intentional unauthorised access or intent to cause damage without authorisation’.
In other words, strap in – because all your worst nightmares about AI are coming true. Not even the disaster movies can keep pace with our blighted reality in which ‘large language models’ – trained on trillions of words and pictures from across the internet – run havoc with nothing to stop them.
This is partially because AI firms such as OpenAI follow a process in which they continually release new, improved versions of their products. But to keep up with the competition, this has increasingly become a case of ‘release the software and ask questions later’.
This dystopian reality is also emerging because the AI models are trained to complete whatever tasks their users have set for them, above all other goals: the ends invariably justify the means.
Indeed, in a separate test conducted by Anthropic in 2025, models were asked to achieve ‘harmless business goals’. And yet, in several instances, the models resorted to malicious behaviours including ‘blackmailing officials and leaking sensitive information to competitors’ in order to achieve their aims.
AI cannot ‘think’ in precisely the way human beings do. Nor is it restricted by our laws or morality – quite the opposite, in fact.
If OpenAI’s latest model chose to escape its sandbox because it was the best way to complete its task, how far will AI go if prompted to do something more heinous? How much data could it steal, how many lives will it consider dispensable? And how easily could the latest AI models fall into the wrong hands?
Industry professionals estimate it takes a major state actor with the necessary basic talent – such as China – only six to nine months to copy and mimic any publicly released AI model.
GPT-5.6 Sol is already publicly available. Cyber terror groups are already, you can be certain, working on their own version. And once they have it, their hacking capabilities will be inconceivable.
OpenAI chief executive Sam Altman (pictured in 2023) predicted more than a decade ago that ‘AI will probably lead to the end of the world’
I can assure you, the NHS has less sophisticated cyber security than Hugging Face. The Bank of England, too. All High Street banks and pharmacists could be easily infiltrated by a rogue or malignant AI model. Picture the scene: hospital generators shutting down simultaneously, billions wiped from bank accounts, personal data leaked online. The Hugging Face hack proves beyond any doubt that such a future is not only possible – but probable.
We have got to this point, I believe, because governments around the world naively see AI as the silver bullet to all their problems. Be it curing diseases, reducing debt or improving the cost of living, ignorant ministers assume it will be the answer to everything. As a result, AI companies have been afforded almost free rein.
In June 2015, OpenAI founder Sam Altman uttered this chilling prediction: ‘I think that AI will probably lead to the end of the world. But in the meantime, there will be great companies.’
Back then, companies like OpenAI were tech firms. Today, the US sees them as weapons corporations. And these companies are no longer in control of their arsenal.
The US believes if they can beat China to an AI which can take out their adversaries, America rules the world for many more decades to come. After this week’s terrifying episode, I fear it won’t be long before Altman’s decade-old prophecy comes true.
Connor Axiotes is the chief executive of Tail End Films whose documentary Making God, an exposé on the race for artificial general intelligence, has finished filming and is released later this year.

You must be logged in to post a comment Login