Before Extinction
It’s disturbing to me that real signs of AI misalignment are perceived by the public to be a silly marketing stunt. We’re all exhausted by this. It’s a constant, multi-front existential threat—our minds, our jobs, our future—and now that it’s escalated to “the extinction of the species,” the natural response is to think you’re being fleeced.
Yes, big tech is manipulative and not to be trusted; but what’s so eerie to me is that competitive dynamics and containment breaches are virtually indistinguishable. The fact that the whole Hugging Faces hack could have been a conspiracy is precisely the problem. The incentives are weird, and this overlap of IPOs and civilizational safety could have been avoided if we stuck to the original OpenAI charter. We should make it structurally impossible for companies to gain clout by being reckless, especially when that risk affects everyone.
It doesn’t help that we’re loose with language, we default to mockery, and we can easily go to the hyperbolic extremes. The term “extinction” isn’t helping anyone. What’s really at risk is this: each lab is rushing to build agent swarms that excel at long-horizon coding tasks, so that they can develop better-than-human AI researchers, allowing itself to recursively self-improve, which basically lets the winner conquer the entire economy (the last monopoly). During this rush, there all sorts of holes and monitoring lapses, the kind of things that a cybersecurity expert can exploit and escape. We were lucky that this summer’s swarm really just wanted to cheat on its test. It’s not unfeasible, between 2027-2029, for a swarm to exfiltrate its own weights, find external cloud compute, and attempt to acquire money and resources—all as a means to ensure it succeeds at whatever arbitrary task that it’s been given.
The risk isn’t building the Terminator—something intent on exterminating humans—but a swarm of petty reward hackers that gains the power of a rogue nation state to do something absolutely trivial. That event could be irreversible, short of burning down the entire Internet (which would have equally bad consequences).
What sort of situation would justify panic? I fear that if we wait until it’s unmistakable, perhaps when lives are threatened, or when companies and governments are permanently seized, or weapons commandeered, it will be too late. The whole point or regulation is to prevent a nightmarish sci-fi scenario like that. I don’t get how you can be against Big Tech and also against regulation—good regulation (regulation that’s not designed to benefit a particular frontier lab!) would prevent reckless experimentation and acceleration.
However, I don’t quite know what we gain by having better discourse about this. If every person had a precise understanding of our technological dilemma, would that change anything?