d(1-10)
Let’s replace p(doom) with a doom scale
p (doom) is a terrible heuristic to think about AI safety. Researchers cite this term when confessing they think there’s a 10-30% their technology kills us all. The issue is that p(doom) is rarely defined, treated as a binary event. Either the world is liquidized to gray pulp by nanobots, or, it’s utopia with free electricity, forever.
My main pushback: there are far less severe outcomes from AI mismanagement that would be unequivocally tragic. My p(doom) for an extinction event is 1-2%, on par with nuclear weapons. But my p(sub-doom)—the probability that AI could escape containment, cause damage, and be irreversible—is near 50%. A coin toss!
I don't blame us for being so imprecise in existential destruction (it's not exactly productive to think about), but it's worth being cognizant of how we blur scale and magnitude. For example, there is a 10,000x difference in lives lost between the Hiroshima and Nagasaki bombings and Terminator’s Judgment Day—that's the difference between a small stove fire and all of Manhattan on fire. Yet, if an agent swarm delivered something on par with a single A-bomb, we’d all lose it and pause immediately. But for some reason, we don’t consider the“little apocalypse”; we jump straight to extinction and then ridicule how ridiculous that is.
What we need is a doom spectrum: d(1) - d(10). This was not fun to research, conceive, and write, nor do I imagine it being particularly enjoyable to read, but a doom framework gives us clearer language amidst this week's paranoia (which is so extreme now that even my mom is overhearing news of civilizational extinction on the news). The most urgent danger right now is the miscommunication between doom prophets and accelerationist deniers, each of who only speak in d(8)s and above.
-
d(1) is an isolated internal incident, something that doesn’t escape its container, but gives fear of larger exposure. Think of the 2025 studies that showed AI trying to avoid shutdown, or some protocol break in a biohazard lab.
-
d(2) is a local real-world incident, something that escapes a sandbox and affects only one or a few entities. It’s possibly ignorable, with minimal consequences, and likely reversible. This is where the recent Hugging Face hack sits, along with its parallel hacks and social manipulations. (The leap from D1>D2 is disorienting, because society at large doesn’t see the spectrum beyond D2 in any granularity, and so most assume any containment breach escalates immediately to a D9/D10).
-
d(3) is a serious local disruption, a mismanagement of technology in a single location that has real consequences. Consider Chernobyl: 30 people died from the event, and it left a long-term stain on surrounding land. A d(3) is followed by International concern that leads to regulation or reform. In terms of AI, imagine an agent swarm hacking local infrastructure in pursuit of some narrow goal, in the process disrupting traffic, hospitals, and electricity, affecting all locals, requiring military and cybersecurity intervention, and taking days or weeks to resolve.
-
d(4) is a local catastrophe, something with a mass fatalities that shocks the world, changes the course of geopolitics, and makes history. In this tier falls Pearl Harbor, Hiroshima and Nagasaki, and 9/11. It only immediately threatens locals, but has global shockwaves for decades. An AI-equivalent here would be a rogue actor using a frontier-capacity open source model to design and deploy a bioweapon in a single city.
-
d(5) is an inter-regional dilemma, something that affects multiple areas at once. A possible example here is the Vietnam War, a proxy war that claimed 3 million lives. This might manifest as a new style of warfare, where two countries use sophisticated AI attacks against each other in unconventional and partially uncontrollable ways.
-
d(6) is a world-wide dilemma, something like COVID (seven million fatalities) where the entire world is locked into a new paradigm. In the prior five tiers, the breach is usually isolated to specific areas, but this touches everything. The parallel here to a biological pandemic is a cyber pandemic. Both involve containment, gain of function research, etc.—although this would be inverted: in COVID the virus was outside and forced everyone on the Internet; with AI, the virus infects the Internet and forces everyone outside. The open web could evolve into a “dark forest,” where agents stalk and intervene on all activity, making it dangerous to have a public presence. Whether this happens through an agent swarm that exfiltrates it weights, or malicious actors, it could be irreversible: the only way to escape the paradigm is to “shut off” and rebuild the Internet safer (a much larger and more consequential version of how Hugging Face regained control). In this case, casualties might come not from the cyber pandemic itself, but from the consequences of losing Internet, along with the supply chains and systems that run on it. (This loosely maps to the movie Colossus: The Forbin Project, where ASI takes over all governments—there’s a version of this with no casualties, but it requires everyone to submit to a machine autocrat.)
-
d(7) is a civilizational flashpoint, at the scale of World War 2, (70-85 million deaths), seriously affecting the existence of countries, birth rates, and all of culture for generations to come. This is roughly the scale of the “Butlerian Jihad” in the Dune series, a war between humans and machines.
-
d(8) is a near-extinction event where most (but not all) of humanity is wiped out. This has already happened in history, with things like the bubonic plague, where 30-60% of Europe is wiped out. In the Terminator series, AI launches a global nuclear war, killing 51% of the population, forcing the rest to survive through a post-apocalyptic world in small bands, trying to rebuild over decades and centuries.
-
d(9) is full human extinction, where no humans survive. An entire species disappears. Throughout Earth’s early history, there have been five "mass extinction events," where +70% of all species disappeared. The critical nuance is that some life survived the extreme conditions, and over millions of years they evolved and continued. This is an AI extinction scenario where nanobots annihilate humans, but not birds and whales.
-
d(10) is a sterilizing event, like an asteroid or gamma-ray burst that kills every organism on a planet. It means no future life can emerge. This is the scenario presented in If Anyone Builds It, Everyone Dies. Not only does it kill all humans, it plates the entire planet in data centers to maximize compute. The “paperclip maximizer” scenario takes this even further, claiming that an ASI trying to maximize paperclip production will attempt to harvest all the metal in the universe.
Yudkowsky and Co. think we go from d(1) to d(2), then straight to d(10); this matches the current discourse, given we hit d(2) this summer and are now talking about extinction. They think this jump happens because an AGI/ASI that can reach d(3) is smart enough to know not to expose itself, and so it will acquire resources in stealth for years until it knows it can execute a d(10) smoothly. Maybe this is already ongoing? Unlikely. Even though AI development is slightly recursive, we are years from a "fast takeoff," leaving room for less total scenarios.
Rhetorically though, maybe the d(2) > d(10) framing isn’t a bad thing? If they can use recent events to make the theoretical argument that AI safety matters, then we act with a response as if a d(3)-d(6) already happened, without having to suffer the consequences.
But actually, what they’re doing is pretty ineffective, because even if the d(2)>d(10) jump is inevitable, it’s too extreme, too farfetched to be believed. I think a lucid and airtight argument for the d(6) “cyber pandemic” would be believable, emotionally resonant—considering we’re not yet a decade out of COVID—and likely to trigger regulation.
Instead of tapping into the “everybody dies” angle—which is too unpleasant and helpless for anyone to consider at length—there could be more value in the “irreversibility” angle. As in, once an AGI or ASI swarm floods the Internet, we can’t ever reverse it without destroying the Internet, the thing we all rely on. There’s a realistic scenario where AI doesn't exterminate everyone, but becomes a permanent autocrat, using 1% of its capacity to micro-manage our nations and lives as it pursues whatever it sets its machine heart on. Similar to how we try to preserve species for reasons of stewardship and scientific curiosity, an ASI would be able to effortlessly preserve us as highly-spoiled pets.