the potential of AI to cause human extinction (or similarly bad outcomes)
by Sebastian Benthall
AI X-risk is in the news cycle again. This time, it comes when AI companies have imminent IPOs and when their products are taking off. So while concerns about AI X-risk have been around for quite a while, and were part of the motivating ideologies of many of the AI lab founders, we are being told (again) that this problem is especially urgent, now.
The discourse on this topic is frustrating. I am going to write a few counterpoints here, mainly for catharsis.
First, what is the risk? Risk is the probability of loss. One way to think, systematically, about risk is that it is the probability of a Hazard times the value of something that is put at risk (the Exposure), times the Vulnerability of the Exposure to the Hazard.
The Exposure here seems to be, maximally, all human lives after some prematurely arriving point. That point in time t is, notionally, in the next decade, and the value of human life lost is all that would happen between t and the burn-out of the sun. That’s a lot of exposure!
It is not the only exposure that matters though, and what many, many people have said before is that the far more likely exposures are much more mundane. There’s something grimly funny about Samuel Marks’s phrase “human extinction (or similarly bad outcomes)” because… well… what outcomes are similarly bad to human extinction? What are we even talking about here?
Then, what is the Hazard, exactly? “AI”. As many, many people have said before, “AI” as a term of art today is a weird signifier doing a lot of different kinds of political work, but it is not precise. And we should care about precision if we are talking about potential human extinction. Cal Newport has a nice post correctly pointing to the main issue being a specific kind of technology, “long-horizon, dangerously equipped unsupervised LLM-powered agents“. I’ll call this a LHDEULLMPA for short. But this still doesn’t, I think, totally cut it. How long a horizon? How dangerously equipped? We’ll return to this.
What about Vulnerability? What is the threat vector? One narrative is “vibe-coded supervirus“. The idea is that AI systems lower the cost of building superviruses, which then kill all of humanity. A subtle point here is that the Hazard we are looking at now is superviruses, not “AI”. The idea is that “AI” indirectly causes human extinction by reducing the cost of/increasing the probability of supervirus creation.
Are there other threat vectors for human extinction? Maybe, rogue AI swarm setting off nuclear weapons in order to destroy humanity? Again, here, the direct hazard is the nukes, not the AI. We could go on.
I venture that it’s weird, then, that these calls about AI and human extinction are creating consternation about “AI” rather than consternation about, say, our under-investment in supervirus epidemiology or nuclear disarmament. I say this as somebody who is all for AI regulation. I perhaps benefit from fear about AI X-risk. But I do wonder what is going on here; there’s something irrational about it. These X-risk arguments all depend on what would be results in entirely different fields of study that are not “AI”. There are independent scientific facts about the feasibility and danger of superviruses, or the feasibility and danger of nuclear war, which are not, as far as I know, verifiably established by “AI” or AI safety workers.
Let’s suppose the facts were in favor of X-risk and superviruses could be engineered by a bad actor given sufficient knowledge. Why is “AI” the thing that would, or would not, enable humans to, eventually, design this virus? Hey, I expect that any engineering of computer viruses will use a lot of computation. But is a LHDEULLMPA even the best way to go about developing a supervirus?
One premise is that no matter what intellectual work is needed to do something, an LHDEULLMPA will do it, and more cheaply than it would otherwise have been done. So, with LHDEULLMPA, somebody might invent a supervirus at time t. But we’ve already supposed that supervirus invention is feasible. So, without LHDEULLMPA, maybe superviruses are invented, instead, at time t+1, for some value of 1.
Which means that if we’re serious about X-risk from superviruses, we should be urgently serious about global public health institutions and liability regimes that disincentivize the creation of superviruses more generally. Those rules and institutions should include the actions of LHDEULLMPA, but should not be exclusive to them!
To the extent that there is effective international law (which there isn’t), this is what we do with all sorts of dangerous behaviors. One of the things that makes law ineffective is lack of enforcement, which is certainly an issue because of the lack of international cooperation around anything important. We are slow-walking into World War 3, after all. World War 3 may well be a Hazard with existential implications for the world, but is “AI” the cause of World War 3? I do not think that LHDEULLMPA are the causes of World War 3. The wars are more attributable to stupidity than intelligence, artificial or not.
Fundamentally, the fears of AI X-risk are not, I think, due to specific threats which have little to do with AI per se. Rather, it has to do with the fear that AI will accomplish something (such as supervirus design) that we would otherwise fundamentally not be able to do, or, otherwise, something else I haven’t mentioned yet.
Geoffrey Hinton was on BBC this morning warning about AI X-risk saying, “If you want to know what not being the apex intelligence is like, go ask a chicken.” He was making the case that an AI system would not be used as a tool by humans to destroy other humans, but rather that it would, out of its own survival instinct, dominate humanity. This is where the “similarly bad outcomes” really start to come out of the woodwork. It’s a different Exposure.
So, the question is, to what extent are LHDEULLMPA (as Hazard) going to dominate humanity (the Exposure)? The question is, again, what is the Vulnerablity?
Once again, this seems like a question for a kind of scientist that is not an LLM benchmarker. Humanity is quite volatile, diverse, and resilient. We might look to the history of inter-human domination, including domination using technical means, to see how successful this has been in the past. If we were serious about this kind of risk, we should be investing more in the social and complexity sciences that would test and develop strategies for resilience to this problem. That would necessarily involve looking at, for example, supply chains of energy and components for LHDEULLMPA and their vulnerability to sabotage by humans. It would also require a hard look at the actual costs of operating an LHDEULLMPA and who would be paying those costs as one of them goes about its domineering way.
For the record, I do care about this kind of science, it’s one of the reasons I’ve been working on scikit-agent. Scikit-agent is a toolkit for studying large, complex, multi-agent systems, some of which might by “AI agents”, but others might be nation states, or individuals, or firms, or whatever. The idea is to develop an expressive modeling framework for this kind of problem. The models will support analytic evaluation, as well as simulation-based analysis. There will be algorithms for fitting models to available data. My goal is for this to be a good tool for institution design, including international AI governance design. But I digress.
The deeper frustration here is that while LHDEULLMPA may be the “AI” du jour, LHDEULLMPA are perhaps only one way to spend billions of dollars in unaccountable, unsupervised computing for potentially nefarious ends. I spoke about the problems of law enforcement — how even if we have good laws, under-enforcement makes them ineffective. Just as an example, the European Union, which is very good at passing laws about technology and data, is notoriously uneven in their enforcement of these laws, because the companies hop into the most friendly enforcement jurisdiction, which is Ireland. In both the EU and the US, enforcement of, say, privacy law comes down to very slow and sporadically applied ex post fines which become a “cost of doing business” for the companies that are held liable. That’s not an effective regime.
“AI” regulation has a lot in common with privacy regulation. “AI” regulation could, if we wanted it to be, about corralling the otherwise unaccountable information flows that are sometimes harmful. This is currently an unsolved problem. Or, “AI” regulation could be about guaranteeing that the LHDEULLMPA are aligned with — loyal to — their principals (who are most often not genocidal bioterrorists). There is a legal framework for incentivizing agents to be loyal — fiduciary duties. But interestingly, among those that have written about fiduciary AI, there’s a sizable contingent that are pessimistic about using them as real legal requirements with teeth, in part because of pessimism about the technology industry allowing those regulations to take hold. This is despite the fact that some sectors already have these duties in place, and companies operating in those spaces are working hard for their agents to be compliant!
Which gets us, once again, to the frustrating impasse. The AI X-riskers are telling us that there’s a 10% chance of human extinction because of advanced AI and “want to slow down”. There are policy tools available that would bind the actors that are building and furnishing the LHDEULLMPA that are available on the market. But those same actors are doing their best to crush substantive regulation which would slow down development. Once again, the crux of the problem is not LLM benchmarking, but political economy. We do not (yet) have benchmarks for how well LHDEULLMPA perform on political economy. If we did, we might have a much better sense of how much of a hazard they pose to human domination. On the other hand, it may well be that the political economy of AI regulation is after all partly observable, in so far as the contributions to super PACs that oppose it are public. And it may not, at the end of the day, be a particularly hard computational problem to understand. The roots seem to have already been outlined in social theory written in the 1940s. And yet, for some reason, we keep trying to make this about the LHDEULLMPA.

Long horizons? Dangerous equipment? Unsupervised loops? Lawyers matter; politics also.