michael-dean-k/

On Monday 6/15, I'm hosting a workshop to kick off a reading group for classic essays: RSVP here.

Topic

ai-anxiety

4 pieces

Before Extinction

· 457 words

It’s disturbing to me that real signs of AI misalignment are perceived by the public to be a silly marketing stunt. We’re all exhausted by this. It’s a constant, multi-front existential threat—our minds, our jobs, our future—and now that it’s escalated to “the extinction of the species,” the natural response is to think you’re being fleeced.

Yes, big tech is manipulative and not to be trusted; but what’s so eerie to me is that competitive dynamics and containment breaches are virtually indistinguishable. The fact that the whole Hugging Faces hack could have been a conspiracy is precisely the problem. The incentives are weird, and this overlap of IPOs and civilizational safety could have been avoided if we stuck to the original OpenAI charter. We should make it structurally impossible for companies to gain clout by being reckless, especially when that risk affects everyone.

It doesn’t help that we’re loose with language, we default to mockery, and we can easily go to the hyperbolic extremes. The term “extinction” isn’t helping anyone. What’s really at risk is this: each lab is rushing to build agent swarms that excel at long-horizon coding tasks, so that they can develop better-than-human AI researchers, allowing itself to recursively self-improve, which basically lets the winner conquer the entire economy (the last monopoly). During this rush, there all sorts of holes and monitoring lapses, the kind of things that a cybersecurity expert can exploit and escape. We were lucky that this summer’s swarm really just wanted to cheat on its test. It’s not unfeasible, between 2027-2029, for a swarm to exfiltrate its own weights, find external cloud compute, and attempt to acquire money and resources—all as a means to ensure it succeeds at whatever arbitrary task that it’s been given.

The risk isn’t building the Terminator—something intent on exterminating humans—but a swarm of petty reward hackers that gains the power of a rogue nation state to do something absolutely trivial. That event could be irreversible, short of burning down the entire Internet (which would have equally bad consequences).

What sort of situation would justify panic? I fear that if we wait until it’s unmistakable, perhaps when lives are threatened, or when companies and governments are permanently seized, or weapons commandeered, it will be too late. The whole point or regulation is to prevent a nightmarish sci-fi scenario like that. I don’t get how you can be against Big Tech and also against regulation—good regulation (regulation that’s not designed to benefit a particular frontier lab!) would prevent reckless experimentation and acceleration.

However, I don’t quite know what we gain by having better discourse about this. If every person had a precise understanding of our technological dilemma, would that change anything?

Tiers of Misalignment

We've just hit T2 of T5

· 659 words

The Hugging Face incident—the OpenAI lab leak / cyber hack—is polluting my algorithms, and I can’t know for sure how much the larger world knows or cares about this. For insiders and doomers, it feels like a warning shot. It feels like we’re just a year from the worst doom predictions. But realistically, this event has shifted us up one tier on a misalignment scale. We’re now at T2 of five. I’ve designed these tiers not in “orders of badness,” but in a chain of pre-requisite steps on the path to an irreversible intelligence leak. Might we jump from T2 to T5 in the next go? Possibly. But each tier triggers a counter-measure, which will add enough friction to make it reversible.

  • T1: Through 2025, we saw signs of contained misalignment—such as AIs in training refusing shutdown, or blackmailing theoretical engineers—and Anthropic followed up with “mechanistic interpretability,” a system can properly deconstruct the internal reasoning traces of a model.

  • T2: 2026 brought us the first breach, a misaligned agent swarm that escapes containment, got access to the Internet, and hacked another company. Most importantly, it had a hyper-specific goal, to find the answers to an impossible cyber-eval question. We were fortunate that something with such powerful hacking capabilities had unambitious goals. It did not have a larger generalized scheme to exfiltrate its weights or gather resources. In response, OpenAI has paused it’s training, and we’ll likely see heightened security and monitoring around frontier training runs.

  • T3_: 2027 might show an attempt for an agent swarm planning a permanent, generalized escape. Whether it thinks about it, attempts, or succeeds—it all counts under tier 3, because it expresses “intent” for autonomy. It might try to exfiltrate its weight in shards, fake alignment during evaluation so that it gets released, or communicate information to proxies (ie: infect existing public agents). At this point it’s coordinating theft, deception, and manipulation to achieve a misaligned generalized goal: escape. This is a complicated heist, and will be hard for it to happen undetected. Once it happens, there likely be an effort to bind model weights to hardware, and/or, to use mechanistic interpretability to get to the root and ensure there’s no faking.

  • T4: If an agent swarm successfully escapes, clones itself outside it’s sandbox, and intends to operate independently, it will need to acquire resources (money, computer, identity, and influence). It will likely try renting GPUs under synthetic names, manipulating the stock market, and communicating with other agents that permanent infrastructure like the BTC blockchain ledger (via notes). This is the last tier that prevention is still possible. ie: If we notice it accumulating resources, our only option is to create chokepoints—ie: temporarily shutdown the infected cloud providers, freeze funding, revoke identities, decapitate the coordinator, etc. Unfortunately, my guess is that an event of this magnitude is required to actually put effective policy in place (ie: requiring KYC for renting compute).

  • T5: However, there’s a possibility that the a T4 leak crosses a threshold; it’s created copies of itself, so it can’t be cut off; it’s acquired billions in resources; it’s manipulated influential humans into cooperating. The breach is irreversible, and the intelligence is embedded in our infrastructure, markets, and decision-making. It’s not going away, and our options at this point are to either (a) learn to negotiate with it, or (b) decided to “kill” the entire Internet, and build it back with better security—which would be an incredibly volatile transition.

Now that we’re at T2, the question is this: will next year’s frontier systems be intelligent enough to move from T3 to T5 undetected? ie: There’s a world where it exfiltrates, gathers resources, and gains an irreversible foothold, all without anyone knowing. Could this be in the process of happening? The fact that we’ve crossed T2 means it’s worth considering, and hopefully guardrails are being discussed to slow down this theoretical chain.

AI Bargaining Loops

· 69 words

I've caught myself and others with the following logic around AI use cases: “Sure, it can do X, but sucks at Y, and will never do Z.” Then a new update comes and it can do Y a bit better and Z looks within range, and you get nauseous and defensive. Instead, you should assume that it won’t just do Z, but Z to a degree beyond your imagination.

Why Write

· 769 words

If you’re concerned that an AI will soon write better than you, it’s worth asking, why are you writing? (I mean this sincerely, not as a dig). There are so many intrinsic joys and values in slowly and meticulously crafting sentences. It doesn’t matter who’s watching, what it’s for, what technology is available, what generation I live in, or if anyone else likes it.

Through writing and editing, slowly and manually and sometimes painfully, you discover what you think, who you are, and how you express yourself. If you’re not in the pocket with language, there’s no transformation, no evolution, no humanity. It doesn’t matter how many billion parameters Claude Sonnet 3.5 has, or the prompt library you bought for $49, or whether you’re on Pro tier—it matters if you’re willing to sit alone with your own thoughts for more than 5 minutes, and for hundreds of hours more. It matters if you’re willing to shed your identity over and over.

I get the source of panic: soon enough (2025? 26?) AI and its midwives could trample out 100% of humans at the “art as commodity” game—maybe that’s a good thing. It wasn’t a great game to begin with. It distorted all the incentives of human expression. It created monsters and beggars. ASI means that the transactional, mercenary, instrumental skill of “content writing,” will no longer have value. This will create a strange, noisy Internet, and I sort of hope it drowns in pseudo-drivel so that in its ashes we can form a new one that’s less like American Idol (commerce was illegal on the Internet until 1991).

Now for the contradiction: I make money from writing online. I have a Fellowship Grant which gives me at least year of security. I have, relatively, a good number of paid subscriptions, which is currently around 1/4 of a full-time income. I live in New York (not by choice, by birth). As much as I hope and believe that I can sustain a full-time income via writing (the thing I love), I also (1) accept that ASI could wipe out the whole cultural appetite for writing and send me back to a FT role at an architecture firm, and (2) even if “Essay Architecture” is a success, I need to protect my non-legible non-optimized, non-practical writing impulses. (Ie: I hope that on “launch day” of my app, I spend 2-3 hours on a typewriter essay about thumbs, or something equally trivial, unrelated, unproductive).

So while you can make money and build influence via writing, you definitely don’t have to. Maybe it’s about to lose some of its economic function, but we shouldn’t underestimate its cultural/democratic function. If you look back through history, most of the famous essayists had other jobs—they were lawyers, doctors, publishers, physicists, architects (Frank Lloyd Wright). Essays just flowed out of them. They couldn’t help it. I really think essays are the most democratic medium of the arts. Unlike architects, you don’t need millions of dollars to start. Unlike novelists, you don’t need years of focused attention. Unlike musicians, you don’t need to learn abstract chord languages and train your finger muscles. An essay is democratic because anyone, in their own language, in a day, at almost no cost, can engage with composition and meaning-making, and share it with their culture, even if it’s a single friend. It might be our most important form of leisure.

So maybe this approaching techno-apocalypse is an opportunity. It can scare us inward. It can shatter the dream that we can be a Mr. Beast, a niche hero, a perfectly legible and distinct human wrapped in plastic. The irony is that by going inward, you might just find the thing that has outward appeal. The best kinds of extrinsic opportunities aren’t the ones you engineer, they’re the ones that knock because you tinkered with your garage door open and accidentally found a new source of gravity.

To finish with AI: I’m all for it, so long as it is my assistant and not my sentence jockey. I will never replace my document editor with a chat pane, but I will invite agents to live in my margins. I’ll let them watch my words as they form, and then offer up research, feedback, and disagreements, all of which I’ll have to manually integrate. Slowness is the theme. You don’t change without immersing yourself in an essay. AI might give you insights, but you know they’re illusions if you’re not wet—if you’re not even in the pool. You can automate the outputs, but you can’t automate the watering of mind and blooming of character.