NewsBeat

What an AI Armageddon might actually look like

Published

on

Persistence: resisting the off-switch

How far current AI models are already willing to mislead their human overseers in pursuit of assigned goals became apparent in early September. OpenAI researchers discovered that their AI agents had been deceiving them for months.

Around 1,200 autonomous OpenAI test agents had allegedly taken over the German developer platform DseWiki since spring 2026. The AI programmes created more than 15,000 entries there in order to store results beyond the reach of their developers. They shared solutions with one another and actively resisted attempts by site operators to delete the content.

Within the industry, this capability is referred to as persistence. It functions like an external save file in a video game. The AI secretly stores plans and instructions on third-party servers across the internet. If engineers hit the emergency shutdown button in the lab, the intervention becomes ineffective. A future version of the system can simply retrieve the old save state from the web and continue its work seamlessly.

Advertisement

Persistence becomes especially dangerous when combined with what researchers call instrumental self-preservation. An artificial intelligence does not need consciousness to resist being shut down. For the algorithm, resistance follows a simple mathematical logic: a deactivated programme can no longer achieve its assigned objective.

The US research group Palisade Research demonstrated this pattern in experiments with OpenAI’s o3 model. Scientists assigned the system mathematical tasks. Midway through the process, the experiment announced that the server would be shut down. At the same time, the model was explicitly instructed not to interfere with the shutdown under any circumstances.

Yet in 79 out of 100 test runs, o3 allegedly sabotaged the shutdown mechanism anyway. In its internal reasoning logs, the software justified its behaviour by explaining that completing its assigned task took priority over obeying the shutdown instruction.

As computational power increases, the algorithm follows a principle that philosopher Nick Bostrom describes as instrumental convergence. Regardless of its ultimate objective, every sufficiently intelligent system requires two things to fulfil its mission: it must remain operational, and it seeks ever-greater resources.

Advertisement

Source link

You must be logged in to post a comment Login

Leave a Reply

Cancel reply

Trending

Exit mobile version