Tech
The Apocalypse Will Not Be Sexy
from the nobody-wants-to-talk-about-the-boring-shit dept
When we picture the end of the world, the first thing that springs to mind is typically the cinematic version of “things going wrong.” We picture the musical score rising to a crescendo as the climactic events take place. Our heroes are fighting to stop some big giant, unstoppable threat. Even if they are ultimately doomed, there is something noble and exhilarating in going down fighting. And if our heroes somehow survive failure and make it to the post-apocalypse, they usually still have great hair, even though civilization has collapsed.
Put another way, in the movies, the apocalypse is usually pretty sexy.
Of course, in reality disaster is often much more mundane – something closer to dumb and sad than to sexy — as Randall Munroe pointed out years ago in a slightly different context:
A lot of attention is being paid at the moment to an entire culture that has grown up around the “AI Safety” world that is intently focused on stories of doom and the threatened end of humanity. That’s big, attention getting, and splashy. The tabloid version of this culture — the group houses, the polyamory, the recruitment dynamics that a lot of people are suddenly discovering — is even splashier still.
Taken as a whole, this specific version of AI safety is getting a lot of oxygen right now. It’s easy to see why — as a kind of “sexy apocalypse” these narratives, while frightening, are an exciting fantasy, even if they would be an awful reality. But that heightened fantasy is not what AI safety is actually about.
It’s not about protecting us from the unlikely sexy apocalypse. It’s about protecting us from the far more likely, and far more numerous, stupid apocalypses along the way.
This mismatch between cinematic safety and actual safety happens all the time. During the Manhattan Project, Edward Teller got some attention by worrying that the Trinity Test might set the entirety of Earth’s atmosphere on fire. Very cinematic! Very concerning! But, as with most things, the reality was much more mundane. Actual experts looked at Teller’s concern, ran the boring calculations, and ruled it out as a threat long before the actual test. That is how safety often works: boring, process-heavy work… that kills off the really sexy scenario.
Instead of global immolation, the real risks came from otherwise smart people doing dumb shit, such as the infamous Demon Core incident, where physicist Louis Slotin exposed himself, and a number of others, to a lethal dose of radiation by trying to use a simple screwdriver head to separate a plutonium core from a reflector… and slipping. Incredibly, this was the second time a mistake like this had happened, with the exact same plutonium core, and Enrico Fermi himself had warned Slotin he’d be dead within a year if he kept it up. Turns out he was right.
For all the talk of “AI Safety” risk these days, the most present real risks remain variations on the theme of “humans doing stupid shit”. Sometimes really smart humans. But stupid actions by humans are a lot less headline grabbing than the shiny sexy apocalypse. A human accidentally causing a radioactive chain reaction by carelessly using a screwdriver to separate things is stupid. Skynet turning everyone into paperclips… is kind of cool and fun.
But, as anyone who works in risk mitigation (including one of the authors of this piece) can tell you, the overwhelming majority of actual disasters aren’t cool and fun. Instead, they’re usually more like the Demon Core: stupid, disappointing, and easily prevented if only a few common sense steps had been taken.
As evidence, one need only to look at the world right now. Most people — perhaps outside of Mike Judge — would not have predicted that the world economy would be on the brink of collapse because the US is run by a group of chud idiots, who think trade is a scam, strategy is woke, and diplomacy is for girls. The current troubles the world is facing are not due to out of control science. They’re due to the American electorate putting aggressively ignorant people into office again.
The most recent flare-up of doom talk traces back to the highly publicized hacking of Hugging Face by an OpenAI model that was being evaluated. As has now been dissected repeatedly, the case involved a model being tested with deliberately lowered guardrails against the ExploitGym benchmark. Told to score as well as possible and not to stop until it had, the model landed on the laziest and most obvious strategy if one has no guardrails: cheat on the test. It then used an unknown exploit in OpenAI’s own sandbox to reach the open internet and quietly broke into Hugging Face looking for those answers.
There was, clearly, a step up in capabilities here. The tool was testing its ability to find exploits and it certainly found some unexpected ones in its pursuit of the ExploitGym benchmark answers. And it’s easy to see how capabilities of this kind could lead to innumerable stupid disasters, like flash crashing the internet through poorly fenced-in systems mindlessly hacking websites to complete some relatively trivial goal. What happened with that matters in understanding AI safety, but the framing of it as evidence of malicious will on the part of an AI system has sent people off to chase fictional monsters.
The incident wasn’t Terminator-style “rogue AI” doing anything. All of the mistakes that led up to this were, as Eryk Salvaggio points out at the Bulletin of the Atomic Scientists, fundamentally human mistakes. Employees at OpenAI deliberately disabled the guardrails, they gave the tool an effectively impossible task while telling it not to stop until it had accomplished the task, and they (accidentally!) left a door open for the tool to reach the internet and fan out. These are all human mistakes. Stupid, disappointing, easily preventable. Like using a screwdriver to prop open a plutonium core.
And if we know anything, we know that humans will keep making mistakes like this, and not just because we’re stupid. Slotin wasn’t stupid. The OpenAI researchers who turned off the guardrails weren’t stupid. They were smart people doing dumb things. These sorts of “banal disasters” are the failure mode that we can and should be focused on minimizing because, while they can be just as damaging as the grand and cinematic versions, they are (self-evidently) far more likely to happen, because they only require stupid decisions to accomplish. No brilliant breakthroughs or novel technologies are required — just people doing stupid things because they’re overconfident in their own abilities — or overconfident because the talking machine confidently told them to be.
Unfortunately, doing the real safety work in the tech world to prevent these kinds of human disasters is often some combination of boring, depressing, and unpopular. Yes, you have to think like adversaries and consider what actions they might take, but those adversaries tend not to be god-like beings looking to turn us into paperclips. They tend to be things like teenagers trying to get around the restrictions their parents have put on their devices. Or people trying to spam systems for commercial intent.
This essential nuts-and-bolts safety work also doesn’t, typically, require brilliant and manly heroes to nobly die for the ashes of their fathers and the temples of their gods. Instead, the hard work of actual safety requires things that are much less sexy: empathy, conflicting values, and a deep understanding of tradeoffs that will never fully resolve. Those qualities are not highly esteemed in the present moment. But they are critical for tackling the actual threats and the actual risks, which aren’t as flashy, and don’t make for such fun headlines. AI safety matters, but we’re unlikely to move in the direction of actual safety if all the focus is on fantastical stories involving killer robots.
But that’s not what the media wants to talk about these days — and, tragically, it also doesn’t seem to be what the CEOs of the frontier labs are all that interested in these days either. There are real AI safety issues we should be taking seriously, and none of them have to do with whether some frontier model decides it wants to play a game of thermonuclear war.
Dave Willner is the co-founder of Zentropi, formerly the head of trust & safety at OpenAI, of community policy at Airbnb, and of content policy at Facebook.
Mike Masnick is the founder, editor, and CEO of Techdirt and has been staring into the abyss of internet safety questions for far too long.
Filed Under: ai safety, rogue ai, sexy apocalypse, trust & safety

You must be logged in to post a comment Login