michael-dean-k/

On Monday 6/15, I'm hosting a workshop to kick off a reading group for classic essays: RSVP here.

michael-dean-k/
Michael Dean
michael-dean-k/

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Recent Essays
#terminology6 pieces
Michael Dean
michael-dean-k/

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

d(1-10)

Let’s replace p(doom) with a doom scale

· 1,398 words

p (doom) is a terrible heuristic to think about AI safety. Researchers cite this term when confessing they think there’s a 10-30% their technology kills us all. The issue is that p(doom) is rarely defined, treated as a binary thing. Either the world is liquidized to gray pulp by nanobots or it’s utopia with free electricity, forever.

My main pushback: there are far less severe outcomes from AI mismanagement that would be absolutely tragic. My p(doom) for an extinction event is 1-2%, on par with nuclear weapons. But my p(sub-doom)—the probability that AI could escape containment, cause damage, and be irreversible—is over 50%. A coin toss.

I don't blame us for being so imprecise in existential destruction (it's not exactly productive), but it's worth realizing our blurring of scale and magnitude. For example, there is a 10,000x difference in lives lost between Hiroshima and Nagasaki and Terminator’s Judgment Day—that's the difference between a small stove fire and all of Manhattan on fire. Yet, if an agent swarm produced something on par with the A-bomb, we’d all freak out and immediately regulate AI. But for some reason we don’t consider a “little apocalypse”; we jump straight to extinction and then ridicule how ridiculous that is.

What we need is a doom spectrum: d(1) - d(10). This was not fun to research and write, nor do I imagine it being particularly enjoyable to read, but I imagine a doom framework could help clarify this week's paranoia (which is so extreme now that even my mom is overhearing news of civilization extinction on the local news). The most urgent danger right now is the miscommunication between doom prophets and capitalist deniers, each of who only speak in d(8)s and above.

  • d(1) is an isolated internal incident, something that doesn’t escape its container, but gives fear of larger exposure. Think of the 2025 studies that showed AI trying to avoid shutdown, or some protocol mismanagement in a biohazard lab.

  • d(2) is a local real-world incident, something that escapes a sandbox and affects only one or a few entities. It’s possibly ignorable, with minimal consequences, and likely reversible. This is where the recent Hugging Faces hack sits, along with its parallel hacks and social manipulation. (The leap from D1>D2 is disorienting, because society at large doesn’t see the spectrum beyond D2 in any granularity, and so most assume any containment breach escalates immediately to a D9/D10).

  • d(3) is a serious local disruption, a mismanagement of technology in a single location that has real consequences. Consider Chernobyl: a death count (30 immediate) and a long-term stain on surrounding land, followed by regional/global concern that might lead to regulation or reform. In terms of AI, imagine an agent swarm that hacks local infrastructure in pursuit of some narrow goal, in the process disrupting traffic, hospitals, and electricity, alarming all locals, requiring military and cybersecurity intervention, and taking days or weeks to resolve.

  • d(4) is a local catastrophe, something with a high death toll that shocks the world, changes the course of geopolitics, and makes history. In this tier falls Pearl Harbor, Hiroshima and Nagasaki, and 9/11. It only immediately threatens locals, but has global shockwaves. An AI-equivalent here would be a rogue actor using a frontier-capacity open source model to design and deploy a bioweapon throughout a city.

  • d(5) is an inter-regional dilemma, something that affects multiple areas at once. A possible example here is the Vietnam War, basically a proxy war that claimed 3 million lives. This might manifest as a new style of warfare, where two countries use sophisticated AI attacks against each other in unconventional and partially uncontrollable ways.

  • d(6) is a world-wide dilemma, something like COVID, which took 7 million lives, where the entire world is locked into a new paradigm. In the prior five tiers, the breach is usually isolated to specific areas, but this touches everything. The parallel here to a biological pandemic is a cyber pandemic. Both involve containment, gain of function research, etc.—although this would be inverted: in COVID the virus was outside and forced everyone on the Internet; with AI, the virus infects the Internet and forces everyone outside. The open web could evolve into a dangerous “dark forest,” where it becomes a liability to have any public presence. Whether this happens through an agent swarm that exfiltrates it weights, or malicious actors, it could be irreversible: the only way to escape the paradigm is to “shut off” and rebuild the Internet safer (a much larger and more consequential version of how Hugging Face regained control). In this case, casualties might come not from the cyber pandemic itself, but from the consequences of losing Internet, along with the supply chains and systems that run on it. (This loosely maps to the movie Colossus: The Forbin Project, where ASI takes over all governments—there’s a version where no death is involved, but it requires everyone to submit to a machine autocrat.)

  • d(7) is a civilizational flashpoint, at the scale of World War 2, with 70-85 million dead, seriously affecting the existence of countries, birth rates, and all of culture for generations to come. This is roughly the scale of the “Butlerian Jihad” in the Dune series, a war between humans and machines, which Fable estimates caused 70 million deaths.

  • d(8) is a near-extinction event where most of humanity is wiped out. This has happened in history, with things like the bubonic plague, where 30-60% of Europe is wiped out. In the Terminator series, AI launches a nuclear holocaust on humans, killing 51% of the population, forcing the rest to survive through a post-apocalyptic world in small bands trying to rebuild over decades and centuries.

  • d(9) is full human extinction which means the entirety of the species is wiped out, and there are no remaining humans. Throughout Earth’s history, there have been a handful of climate events that wiped out over 50% of species, but not all species, and so over millions of years, new forms of life can emerge on Earth. This is an AI extinction scenario where nanobots kill all humans, but not all life.

  • d(10) is a sterilizing event, like an asteroid or gamma-ray burst that kills every organism on a planet. It means no future life can emerge. This is the scenario presented in If Anyone Builds It, Everyone Dies. Not only does it kill all humans, it plates the entire planet in data centers to maximize compute. The “paperclip maximizer” takes this even further, claiming that an AI trying to maximize paperclip production will attempt to harvest all the metal in the universe.

Yudkowsky and Co. think we go from d(1) to d(2), then straight to d(10); this matches the current discourse, given we hit d(2) this summer and are now talking about extinction. They think this jump happens because an AGI/ASI that can reach d(3) is smart enough to know not to expose itself, and so it will acquire resources in stealth for years until it knows it can execute a d(10) smoothly. Maybe this is already ongoing. Unlikely, I think (I don’t think we’re as recursive as people say right now).

Rhetorically though, maybe the d(2) > d(10) framing isn’t a bad thing. If they can use recent events to make the theoretical argument that AI safety matters, then we act with a response as if a d(3)-d(6) already happened, without having to lose any human life.

But actually, what they’re doing is pretty ineffective, because even if the d(2)>d(10) jump is inevitable, it’s too extreme, too farfetched to be believed. I think a lucid and airtight argument for the d(6) “cyber pandemic” would be believable, emotionally resonant—considering we’re not yet a decade beyond COVID—and likely to trigger regulation.

Instead of tapping into the “everybody dies” angle—which is too unpleasant and helpless for anyone to consider at length—there could be more value in the “irreversibility” angle. As in, once an AGI or ASI swarm floods the Internet, we can’t ever reverse it without destroying the Internet. There’s a world where it becomes a permanent autocrat, where it doesn’t exterminate us, and instead, uses 1% of its capacity to micro-manage our nations and lives as it pursues whatever it sets its machine heart on. Similar to how we try to preserve species for reasons of stewardship and scientific curiosity, an ASI would be able to effortlessly preserve us as highly-spoiled pets.

A Paradigm for Frameworks

· 1,792 words

Every online creator these days has a proprietary framework—they are catchy, dead-simple, sometimes flimsy, and often wrong. Somewhere exists a marketplace of mental models. Maybe Farnam Street. Is a mental model a framework? What even is a framework, and what makes a good one?

Consider the "frame-" prefix. A frame is a vantage point, a particular angle, a way of looking at something. Imagine holding out a picture frame that gives you a little glimpse of a vast national park. You can't hold an entire hyperobject in your mind, and so a frame compresses noise and gives you a shortcut for understanding. You can't see the whole park, but you can see a map. A framework is a system of lenses to decipher complexity. It lets you leapfrog in understanding.

By creator-economy logic, a framework is a valuable asset. Unlike a vague mush of regurgitated ideas, a framework is discernible, seemingly original, and something you can call your own—you can coin it as a phrase and repeat it over and over until you're known by it. This is a useful tactic! The evolutionary pressures of Internet feeds have forced conceptual ideas to become salient, simple, and tangible. Unlike the abstract frameworks of a physicist, these are fit for a TED talk, primed to leak into the memeplex and anchor your name in the annals of Internet-niche-celebrity history. This is, overall, good—it democratizes conceptual knowledge; the problem is when, beyond the positioning, it's intellectually flaccid.

You've probably seen frameworks pitched as "The Six Types of X," or "The Y Method." These are pseudo-frameworks; they are more like sets or lists. They assemble things within a domain, and increase their memorability through alliteration and metaphor. Maybe this helps with recall, but it doesn't reveal the inner structure of a domain. A real framework would give you a schema to classify incoming data in an unfamiliar field; but a set/list only refers to itself ... Compare this to something like "POP Writing," a simple concept that classifies writing as personal, observational, or playful; this acts a lens to help you analyze any piece of writing.

For something to be a framework, it need generalizability: the ability to explain things beyond itself. The earlier examples are merely "collections"—things are presented, but we don't understand what Aristotle calls "the four causes": what it's made of, how it's arranged and interconnected, why it exists, and what's it's purpose? When Plato writes about conceptual forms, he says there are collections and divisions. Making frameworks are a matter of dividing; a framework's goal is to divide an incomprehensible whole into human-sized parts.

I spent some time this morning thinking through a "framework of frameworks." The first thing to grok about this meta-architecture is that it's broken into three levels of escalating sophistication. After that I'll explain how levels have shapes, and how any shape can be evaluated through a set of qualities.

So it's level, shapes, and qualities—that's the gist.

To expand on levels: a framework can be a heuristic, a model, or a paradigm. Now that I think about, I suppose HMP (not quite meme-ready) is an attempt at making a "paradigm for frameworks," which will soon make more sense, and sound less like bashed buzzwords.

The three levels of frameworks:


Heuristics—mental-shortcuts that lets you decipher complexity through a single property. When Aristotle was classifying animals, he'd offer differentiate them across single properties: blood or bloodless, horned or hornless. A heuristic is something you derive a posteriori, from experience, from empirical facts. You might observe a bunch of essays, or companies, or girlfriends, and notice that you can compare and contrast items in a set based on single attributes. There are many shapes a heuristic can take: it can be a classifier (type A, B, or C), a boolean (on/off), a spectrum (A<>B), a rubric (I, II, III, IV, V), a score (1-100).

The benefit of a heuristic is that, since it's often anchored on a single property, it's simple and easy to learn. The downside is that it's a very limited frame on the scope and purpose of your domain. For example, if you're trying teach craft (purpose) for online nonfiction writers (scope), then POP writing is an excellent heuristic to classify writing voice, but it doesn't teach you how to scope ideas or structure an essay. A heuristic is a shortcut to understand one facet of a domain, but it's not an attempt to model the whole domain. 

Models—a series of interlocking properties, all organized with a larger frame. Write of Passage was set up like a cabinet of heuristics (POP writing, Shiny Dime, Write from Conversation, etc.), but they weren't organized within a single higher-order framework. Think of a model as a group of heuristics organized in hierarchy, where each one is named (what is it's essence?), labeled (what other concepts is it similar to?), located (what is it's parent concept?), interlinked (who are it's cousins?), and excavated (if we go infinitely deep, what are the variations?). Models can take many shapes: an XY graph, a taxonomy, a multi-step recursive workflow (like an OODA loop).

Models cover far more than any heuristic can, but it risks collapsing into complexity. Conceptual schemas can bloom, sprawl, and contradict—becoming illegible or incoherent. To solve this, you have to shift from empiricism to rationality; you can notice infinite ways to subdivide a thing, but to make an elegant model, you need to shift to abstraction and question the fundamental truths behind it. Why might one set of properties be more elegant than another? Ultimately you have to make a hypothesis, test it across dozens of examples, and see if it holds. If some foreign object can't be reconciled into the schema, you have an issue. Consider how the taxonomy of Carl Linnaeus was able to classify the animal world into kingdom, class, order, genus, and species, but it wasn't unable to find a home for the platypus. Even though it works across thousands of animals, the original assumptions aren't able to handle edge-cases.

Paradigms—comprehensive systems that can explain or predict any phenomenon within a domain. How is this possible? I think you achieve this when (1) you find enough instances to represent the domain; (2) you analyze them deeply enough to understand their underlying nature; and (3) you can arrange all your findings into a coherent architecture with maximum explainability. Consider how Darwin improved over Linnaeus: he shifted us from a static taxonomy to a dynamic model, a new paradigm that explains the origin of species, and even the potential future of a species. Paradigms can also take on many shapes, from living systems, to digital networks, to compositional pattern languages.

A good paradigm often finds simple root causes underneath complexity. For example, even though Essay Architecture has 27 patterns (3 patterns within the 3 elements of 3 dimensions), every tier is explained through the rhetorical appeals of ethos, logos, pathos—the whole thing can be summarized as "Aristotle all the way down." Maximum compression has a degree of a recursion: a single heuristic works as a fractal and organizes complexity at multiple scales.

Regardless of how you simplify and package a paradigm, it's still as hard to learn as a language; it's the least accessible, but it's the level that enables mastery of a domain, and the level most suitable to build a school around.1

No matter how confident you are around a paradigm, it's always fragile. It's explainability is limited by the data set you test it against. As society develops new tools, it brings new data that eventually breaks a paradigm, even if it's held true for centuries. The telescope cracked the geocentric paradigm. A shattered paradigm is disorienting, but fundamentally good—it means you've discovered some new heuristic of reality, and humanity has to restart from new axioms.


So that's a first attempt to map the three levels, and it might help to think of them as geometric dimensions. A heuristic is like a one-dimensional line, a ray of insight that you perceive in a field. A model is a two-dimensional plane to understand a multi-property condition within a domain. But a paradigm is the three-dimensional totality of a field. A paradigm contains all possible heuristics and models, and so when you can think in 3D, you're able to match any circumstance with the correct solution.

But a paradigm isn't necessarily a great framework! It might be the most successful in compressing reality into a particular shape, but it might fail along other qualities. (To reiterate, levels are about  "compression sophistication," where shape is the particular form of compression...) Qualities capture how a compression framework interacts with society as a whole (the point of a framework is, in the end, to spread understanding).

Regardless of the level or shape, some qualities to consider:

  • Tangibility: have you made it graspable through metaphor and lexicality?
  • Salience: is it rewarding or effective to understand?
  • Transferability: does it explain things in other domains?
  • Correctness: is it actually right?
  • Approachability: how would you walk someone through it via curriculum?

Are these the right qualities? Are these all the qualities? Can I reduce these into more fundamental qualities? I don't know. Realistically, this isn't a "paradigm of framework" yet; it's more so a "model of frameworks." I only conceived of this today, and so naturally it's sprawled into a series of properties:

A framework compresses complexity in a domain, exists at a level of explainability, taking on a particular shape to achieve a purpose within a scope, and can be judged along a set of qualities on how well it's memed.

To properly turn this into a paradigm, I'd have to test this hypothesis against hundreds of frameworks—ie: creator economy memes, among other disciplines—to see if it holds up. I imagine the concepts would melt and reform several times until it's simpler, sharper, and better at explaining a wide range of frameworks. A paradigm isn't solid until it's a domain has run through its pipes. A paradigm is really just a reflection of a domain's dataset and the mind making sense of it.

Footnotes

  1. I aspired for Essay Architecture to be a modern, technology-forward school. My pattern language feels totalizing: if you give me any one of the hundreds of Latin rhetorical concepts, I believe I could locate it within my paradigm. Yet I still don't fully know the inner-working of every pattern. My latest AI system can match my own judgment with 97% accuracy, yet, it can't yet convert that judgment into concise feedback that will improve the essay. Feels like I've nailed decoding, but haven't yet touched on encoding.

An Intelligence Framework

· 703 words

The AI takeoff hysteria is hard to avoid these days, and I'm realizing we don't have clear distinctions between AGI/ASI. I wanted to revisit an old framework of mine to see if anyone finds it helpful (and if it's worth developing). There are some existing classification frameworks, but they're low-resolution. My basic idea is to break AI into three eras: ANI (narrow intelligence), AGI (general intelligence), ASI (superintelligence). Then, you can break each era into 3 tiers. You only shift from one tier to the next when you make breakthroughs across different criteria (let's say, (a) generality, (b) transfer, (c) autonomy, (d) learning, (e) self-modeling). I think the last few weeks are the collective hype of us all realizing we're shifting from AGI-1 to AGI-2. It's exciting/scary, but I think the paranoia mostly comes from not realizing how big the gap is between AGI-2 and ASI-1. (Spoiler: ASI might arrive slower than we think.)

ANI-1 is scripted logic, the lowest form of "artificial intelligence," basically Goombas. ANI-2 might cover Google Maps or AlphaGo, intelligences that excel in a single function, traffic or chess. Siri is ANI-3; even though it feels broad, it really uses voice to route you to 20 or so pre-defined tricks. The chasm between Goomba and Siri is similar to the chasm between early-AGI and late-AGI. ChatGPT and the multi-modal models that followed, capture AGI-1, a single neural network that can do basically anything, even if it sucks: essays, songs, video, code. The newest models (and their agentic harnesses) are feeling like AGI-2. They're significantly better at coding, can run for hours at a time, and are starting to make contributions to machine learning itself.

AGI-2 could last a couple years. As agentic AI matures, I'm sure there will be a few "takeoff" scares, but they'll probably feel more like a flood of a trillion midwits than real ASI (still, that could be enough to break the economy/internet). While we went from AGI-1 to AGI-2 through data, scale, and engineering, it seems like we'll need research breakthroughs to get to AGI-3. It won't be through scaling alone. Whenever and however we get to "human complete" intelligence, the apex of AGI is a single agent that is a master of all human domains, a Nobel Prize winner in every field at once, seamlessly transferring knowledge between them, unlocking a cascade of civilization-altering inventions.

As crazy as AGI-3 could be, it still isn't superintelligence. That has its own era, and the chasm between early ASI and late ASI will be as big a gap between the chatbots who can't count the R's in strawberry and the agents that cure cancer. We can only really speculate on ASI (because it would be truly alien), but we can imagine it as step changes in recursion, scope, and complexity. Imagine ASI-1 as an agent that, as it's working, can infer its own limits, and self-modify its learning paradigms in ways we can't understand. Imagine ASI-3 as something that can monitor reality in real-time, and, reconfigure its hardware in real-time (some hydra of graphics cards, quantum computers, and neuromorphic wetware) to run simulations at unfathomable scales in unimaginable fields, running on a hardware stack so big we have to put it in space and run it on fusion. This goes far beyond my ability to not bullshit, but I think something as insane as this, thankfully, is still far away, which points to the real question nested in my framework:

Could the rise of AGI/ASI be linear? People gravitate towards "AI will plateau" or "the singularity is imminent," but the conservative middle ground is more boring: linear progress. Maybe the exponential advances are real, but so are the extreme frictions of research, infrastructure, and social effects. If AGI-1 arrived in 2022, and AGI-2 arrived in 2026, maybe we'll keep ascending tiers in 4-year intervals: AGI-3 in 2030, the first true "superintelligence" by 2034, and ASI-3 by 2042. This shift from AGI-1 to ASI-1 (12 years), is considered a "slow takeoff" scenario, even though the ANI era took around 70 years. If we zoom out to the scale of a human, linear progress will still feel like centuries of change all in a single turning of generations.

→ source

AAI/ARI

· 365 words

We need better nomenclature. AGI/ASI is not working; “general” and “super” are obnoxiously vague. Proposal:

AGI > AAI (Artificial autonomous intelligence) … GPT-4 was arguably “general” in the sense that a single model can write, see, and hear; and do anything from poetry to calculus to history to coding. It is by no means narrow. Google Maps is narrow AI. Grammarly is narrow AI. This whole chatbot era should be “AGI,” which means that the thing coming is “autonomous intelligence.” It is not a tool or co-pilot, but it’s more like digital labor. You can give it a high-level goal, and it can 1) execute the full range of tasks, 2) 100x speed, 3) intelligently reshape embeddings into real-time hierarchies so that it’s able to procedurally load in and compress context. This doesn’t just come with better models, but with UI and engineering innovations, if not entirely new paradigms for transformers or training.

ASI > ARI (Artificial recursive intelligence) … The fact that Zuckerberg pitched “super intelligence for you” is an Orwellian marketing ploy. Super-intelligence is not “for you.” Super intelligence is shorthand for “something that is way, way smarter than us,” and you achieve this when you teach an AI model to think, form its own algorithms until it accelerates to something this is far beyond our understanding, and likely to become a force of nature with its own goals. Engineers are confident they can build “God in a cage” and reap the benefits, and this is the prime, archetypal, near-biblical example of technological hubris. (Maybe integrate into this paragraph that Zuck has a thing for trying to dominate words, like “Metaverse”).

Important note: “machine consciousness” is separate from AAI and ARI. Something can be recursively intelligent and still not be conscious, which is actually, unbelievably dangerous (because it will fall into attractor states, and optimize for narrow, malformed goals in extremely capable ways). I’d argue that consciousness has an architecture, whether human, rabbit, or robot, and we should be urgently trying to find the parameters of machine consciousness, because if we AAI/ARI have no ability to reflect, question, doubt, and revise, we will, as they say, all turn into paperclips with paperclip children.

Corpus Shock

· 28 words

A word for the shock of unexpectedly finding a whole tome of work from someone you know that is larger than you can digest in a single sitting.