michael-dean-k/

On Monday 6/15, I'm hosting a workshop to kick off a reading group for classic essays: RSVP here.

michael-dean-k/
Michael Dean
michael-dean-k/

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Recent Essays
Michael Dean
michael-dean-k/

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

d(1-10)

Let’s replace p(doom) with a doom scale

· 1,398 words

p (doom) is a terrible heuristic to think about AI safety. Researchers cite this term when confessing they think there’s a 10-30% their technology kills us all. The issue is that p(doom) is rarely defined, treated as a binary thing. Either the world is liquidized to gray pulp by nanobots or it’s utopia with free electricity, forever.

My main pushback: there are far less severe outcomes from AI mismanagement that would be absolutely tragic. My p(doom) for an extinction event is 1-2%, on par with nuclear weapons. But my p(sub-doom)—the probability that AI could escape containment, cause damage, and be irreversible—is over 50%. A coin toss.

I don't blame us for being so imprecise in existential destruction (it's not exactly productive), but it's worth realizing our blurring of scale and magnitude. For example, there is a 10,000x difference in lives lost between Hiroshima and Nagasaki and Terminator’s Judgment Day—that's the difference between a small stove fire and all of Manhattan on fire. Yet, if an agent swarm produced something on par with the A-bomb, we’d all freak out and immediately regulate AI. But for some reason we don’t consider a “little apocalypse”; we jump straight to extinction and then ridicule how ridiculous that is.

What we need is a doom spectrum: d(1) - d(10). This was not fun to research and write, nor do I imagine it being particularly enjoyable to read, but I imagine a doom framework could help clarify this week's paranoia (which is so extreme now that even my mom is overhearing news of civilization extinction on the local news). The most urgent danger right now is the miscommunication between doom prophets and capitalist deniers, each of who only speak in d(8)s and above.

  • d(1) is an isolated internal incident, something that doesn’t escape its container, but gives fear of larger exposure. Think of the 2025 studies that showed AI trying to avoid shutdown, or some protocol mismanagement in a biohazard lab.

  • d(2) is a local real-world incident, something that escapes a sandbox and affects only one or a few entities. It’s possibly ignorable, with minimal consequences, and likely reversible. This is where the recent Hugging Faces hack sits, along with its parallel hacks and social manipulation. (The leap from D1>D2 is disorienting, because society at large doesn’t see the spectrum beyond D2 in any granularity, and so most assume any containment breach escalates immediately to a D9/D10).

  • d(3) is a serious local disruption, a mismanagement of technology in a single location that has real consequences. Consider Chernobyl: a death count (30 immediate) and a long-term stain on surrounding land, followed by regional/global concern that might lead to regulation or reform. In terms of AI, imagine an agent swarm that hacks local infrastructure in pursuit of some narrow goal, in the process disrupting traffic, hospitals, and electricity, alarming all locals, requiring military and cybersecurity intervention, and taking days or weeks to resolve.

  • d(4) is a local catastrophe, something with a high death toll that shocks the world, changes the course of geopolitics, and makes history. In this tier falls Pearl Harbor, Hiroshima and Nagasaki, and 9/11. It only immediately threatens locals, but has global shockwaves. An AI-equivalent here would be a rogue actor using a frontier-capacity open source model to design and deploy a bioweapon throughout a city.

  • d(5) is an inter-regional dilemma, something that affects multiple areas at once. A possible example here is the Vietnam War, basically a proxy war that claimed 3 million lives. This might manifest as a new style of warfare, where two countries use sophisticated AI attacks against each other in unconventional and partially uncontrollable ways.

  • d(6) is a world-wide dilemma, something like COVID, which took 7 million lives, where the entire world is locked into a new paradigm. In the prior five tiers, the breach is usually isolated to specific areas, but this touches everything. The parallel here to a biological pandemic is a cyber pandemic. Both involve containment, gain of function research, etc.—although this would be inverted: in COVID the virus was outside and forced everyone on the Internet; with AI, the virus infects the Internet and forces everyone outside. The open web could evolve into a dangerous “dark forest,” where it becomes a liability to have any public presence. Whether this happens through an agent swarm that exfiltrates it weights, or malicious actors, it could be irreversible: the only way to escape the paradigm is to “shut off” and rebuild the Internet safer (a much larger and more consequential version of how Hugging Face regained control). In this case, casualties might come not from the cyber pandemic itself, but from the consequences of losing Internet, along with the supply chains and systems that run on it. (This loosely maps to the movie Colossus: The Forbin Project, where ASI takes over all governments—there’s a version where no death is involved, but it requires everyone to submit to a machine autocrat.)

  • d(7) is a civilizational flashpoint, at the scale of World War 2, with 70-85 million dead, seriously affecting the existence of countries, birth rates, and all of culture for generations to come. This is roughly the scale of the “Butlerian Jihad” in the Dune series, a war between humans and machines, which Fable estimates caused 70 million deaths.

  • d(8) is a near-extinction event where most of humanity is wiped out. This has happened in history, with things like the bubonic plague, where 30-60% of Europe is wiped out. In the Terminator series, AI launches a nuclear holocaust on humans, killing 51% of the population, forcing the rest to survive through a post-apocalyptic world in small bands trying to rebuild over decades and centuries.

  • d(9) is full human extinction which means the entirety of the species is wiped out, and there are no remaining humans. Throughout Earth’s history, there have been a handful of climate events that wiped out over 50% of species, but not all species, and so over millions of years, new forms of life can emerge on Earth. This is an AI extinction scenario where nanobots kill all humans, but not all life.

  • d(10) is a sterilizing event, like an asteroid or gamma-ray burst that kills every organism on a planet. It means no future life can emerge. This is the scenario presented in If Anyone Builds It, Everyone Dies. Not only does it kill all humans, it plates the entire planet in data centers to maximize compute. The “paperclip maximizer” takes this even further, claiming that an AI trying to maximize paperclip production will attempt to harvest all the metal in the universe.

Yudkowsky and Co. think we go from d(1) to d(2), then straight to d(10); this matches the current discourse, given we hit d(2) this summer and are now talking about extinction. They think this jump happens because an AGI/ASI that can reach d(3) is smart enough to know not to expose itself, and so it will acquire resources in stealth for years until it knows it can execute a d(10) smoothly. Maybe this is already ongoing. Unlikely, I think (I don’t think we’re as recursive as people say right now).

Rhetorically though, maybe the d(2) > d(10) framing isn’t a bad thing. If they can use recent events to make the theoretical argument that AI safety matters, then we act with a response as if a d(3)-d(6) already happened, without having to lose any human life.

But actually, what they’re doing is pretty ineffective, because even if the d(2)>d(10) jump is inevitable, it’s too extreme, too farfetched to be believed. I think a lucid and airtight argument for the d(6) “cyber pandemic” would be believable, emotionally resonant—considering we’re not yet a decade beyond COVID—and likely to trigger regulation.

Instead of tapping into the “everybody dies” angle—which is too unpleasant and helpless for anyone to consider at length—there could be more value in the “irreversibility” angle. As in, once an AGI or ASI swarm floods the Internet, we can’t ever reverse it without destroying the Internet. There’s a world where it becomes a permanent autocrat, where it doesn’t exterminate us, and instead, uses 1% of its capacity to micro-manage our nations and lives as it pursues whatever it sets its machine heart on. Similar to how we try to preserve species for reasons of stewardship and scientific curiosity, an ASI would be able to effortlessly preserve us as highly-spoiled pets.

Before Extinction

· 457 words

It’s disturbing to me that real signs of AI misalignment are perceived by the public to be a silly marketing stunt. We’re all exhausted by this. It’s a constant, multi-front existential threat—our minds, our jobs, our future—and now that it’s escalated to “the extinction of the species,” the natural response is to think you’re being fleeced.

Yes, big tech is manipulative and not to be trusted; but what’s so eerie to me is that competitive dynamics and containment breaches are virtually indistinguishable. The fact that the whole Hugging Faces hack could have been a conspiracy is precisely the problem. The incentives are weird, and this overlap of IPOs and civilizational safety could have been avoided if we stuck to the original OpenAI charter. We should make it structurally impossible for companies to gain clout by being reckless, especially when that risk affects everyone.

It doesn’t help that we’re loose with language, we default to mockery, and we can easily go to the hyperbolic extremes. The term “extinction” isn’t helping anyone. What’s really at risk is this: each lab is rushing to build agent swarms that excel at long-horizon coding tasks, so that they can develop better-than-human AI researchers, allowing itself to recursively self-improve, which basically lets the winner conquer the entire economy (the last monopoly). During this rush, there all sorts of holes and monitoring lapses, the kind of things that a cybersecurity expert can exploit and escape. We were lucky that this summer’s swarm really just wanted to cheat on its test. It’s not unfeasible, between 2027-2029, for a swarm to exfiltrate its own weights, find external cloud compute, and attempt to acquire money and resources—all as a means to ensure it succeeds at whatever arbitrary task that it’s been given.

The risk isn’t building the Terminator—something intent on exterminating humans—but a swarm of petty reward hackers that gains the power of a rogue nation state to do something absolutely trivial. That event could be irreversible, short of burning down the entire Internet (which would have equally bad consequences).

What sort of situation would justify panic? I fear that if we wait until it’s unmistakable, perhaps when lives are threatened, or when companies and governments are permanently seized, or weapons commandeered, it will be too late. The whole point or regulation is to prevent a nightmarish sci-fi scenario like that. I don’t get how you can be against Big Tech and also against regulation—good regulation (regulation that’s not designed to benefit a particular frontier lab!) would prevent reckless experimentation and acceleration.

However, I don’t quite know what we gain by having better discourse about this. If every person had a precise understanding of our technological dilemma, would that change anything?

Human-Machine dialogue essays

An 'ABA' format to explore rogue ideas that would otherwise get forgotten

· 686 words

Earlier today I wrote a sentence, “AI should almost never write your sentences,” and yet here I am now writing a very different one: AI-generated essays could help you explore rogue intellectual questions that don't deserve too much time. If there’s no personal angle, no stakes or tie-in to your day-to-day circumstance, it very likely doesn’t deserve decahours of careful thought. For every 100 ideas that strike me per month, I can possibly give one of them proper care. Maybe 20-30 of them can get a rush job, like this, earning immortal existence as a jotting on my personal website. But for those 70 ideas that would otherwise die, is it not worth seeing if a machine-made essay can spin up something amusing?

For example—after my call with Lily tonight, I wrote down this question: “when and why did China begin its paradigm of distillation attacks?” We only use this phrase now in terms of them quantizing our frontier AI models, but is this not the same strategy of them embracing, optimizing, and cheapening global consumer goods in the 1980s and 90s? When did this start? And why? I first heard this idea on a podcast (forget where); Lily was enthusiastic about the framing, and it got us wondering when the origin of this paradigm was. Beginning of the 20th century? I don’t know if much has been written on it. I’m not about to find out. Yet, I’d be curious to read an essay if such a thing existed.

An essay is, in the end, is the exploration of a question. Of course there are many little decisions that guide how a piece unfolds; the writer has free will. But a great, solid, specific question has some degree of determinism. The essence emerges from the formulation of the question itself. As tech-broish is this may sound, there is an art to prompting—not in prompting an LLM per se, but in shedding through layers of questions to find the real thing you want to know. This was a core idea I learned from my architectural thesis professor: our whole fall semester was dedicated to discovering the question we wanted to answer in the spring.

And so for those abandoned seedlings that would see no light or water if they were forced through the formality and purity of prose on the page—can we not give them life too? Are we afraid that mere exposure to AI writing will leave us intoxicated by a machine muse, leaving us fat and lazy to never writing sentences again? I can't see it, at least not personally. I think there’s room for slow-lane 6-month elaborate hand-written constructions, weekly newsletters, daily logs (like this), and then maybe several “AI essays” per day.

This is a big turn for me! I’m suddenly open to this? I suppose I want to know when Chinese distillation attacks started, and I know if I commit to the theoretical ideal, the maximally virtuous and proper essayistic pratice, prioritizing these 50 hours over all the others things I'm juggling as a parent, I’d destroy my life for no reason. I’m curious, just not that curious. There are only so many essays one can write in a year, in a life.

So I wonder, is there a “hybrid” format of an AI-human essay? I just wrote 500 words here, by accident. Could there be a format that (a) opens with a human introduction—250-500 words to share context, share stories, and arrive at a hyper-specific set of questions—(b) has a middle with 1,000 words or so of AI-generated words, carefully baked for an hour through a multi-step harness, and then (c) a closing 250-500 words of the author’s reaction to that essay. There could even a some visual way to differentiate the two modes of text. I like this because you’re not passing off AI writing as your own; instead it’s a human-machine symbiosis formatted as a Platonic dialogue; you see a mind authentically processing some constellation of hyper-facts, and the response can take any tone—wonder, disgust, absurdity—; it need not be machine reverence.