Digifesto

Category: artificial intelligence

the potential of AI to cause human extinction (or similarly bad outcomes)

AI X-risk is in the news cycle again. This time, it comes when AI companies have imminent IPOs and when their products are taking off. So while concerns about AI X-risk have been around for quite a while, and were part of the motivating ideologies of many of the AI lab founders, we are being told (again) that this problem is especially urgent, now.

The discourse on this topic is frustrating. I am going to write a few counterpoints here, mainly for catharsis.

First, what is the risk? Risk is the probability of loss. One way to think, systematically, about risk is that it is the probability of a Hazard times the value of something that is put at risk (the Exposure), times the Vulnerability of the Exposure to the Hazard.


R=H×E×VR = H \times E \times V


The Exposure here seems to be, maximally, all human lives after some prematurely arriving point. That point in time t is, notionally, in the next decade, and the value of human life lost is all that would happen between t and the burn-out of the sun. That’s a lot of exposure!

It is not the only exposure that matters though, and what many, many people have said before is that the far more likely exposures are much more mundane. There’s something grimly funny about Samuel Marks’s phrase “human extinction (or similarly bad outcomes)” because… well… what outcomes are similarly bad to human extinction? What are we even talking about here?

Then, what is the Hazard, exactly? “AI”. As many, many people have said before, “AI” as a term of art today is a weird signifier doing a lot of different kinds of political work, but it is not precise. And we should care about precision if we are talking about potential human extinction. Cal Newport has a nice post correctly pointing to the main issue being a specific kind of technology, “long-horizon, dangerously equipped unsupervised LLM-powered agents“. I’ll call this a LHDEULLMPA for short. But this still doesn’t, I think, totally cut it. How long a horizon? How dangerously equipped? We’ll return to this.

What about Vulnerability? What is the threat vector? One narrative is “vibe-coded supervirus“. The idea is that AI systems lower the cost of building superviruses, which then kill all of humanity. A subtle point here is that the Hazard we are looking at now is superviruses, not “AI”. The idea is that “AI” indirectly causes human extinction by reducing the cost of/increasing the probability of supervirus creation.

Are there other threat vectors for human extinction? Maybe, rogue AI swarm setting off nuclear weapons in order to destroy humanity? Again, here, the direct hazard is the nukes, not the AI. We could go on.

I venture that it’s weird, then, that these calls about AI and human extinction are creating consternation about “AI” rather than consternation about, say, our under-investment in supervirus epidemiology or nuclear disarmament. I say this as somebody who is all for AI regulation. I perhaps benefit from fear about AI X-risk. But I do wonder what is going on here; there’s something irrational about it. These X-risk arguments all depend on what would be results in entirely different fields of study that are not “AI”. There are independent scientific facts about the feasibility and danger of superviruses, or the feasibility and danger of nuclear war, which are not, as far as I know, verifiably established by “AI” or AI safety workers.

Let’s suppose the facts were in favor of X-risk and superviruses could be engineered by a bad actor given sufficient knowledge. Why is “AI” the thing that would, or would not, enable humans to, eventually, design this virus? Hey, I expect that any engineering of computer viruses will use a lot of computation. But is a LHDEULLMPA even the best way to go about developing a supervirus?

One premise is that no matter what intellectual work is needed to do something, an LHDEULLMPA will do it, and more cheaply than it would otherwise have been done. So, with LHDEULLMPA, somebody might invent a supervirus at time t. But we’ve already supposed that supervirus invention is feasible. So, without LHDEULLMPA, maybe superviruses are invented, instead, at time t+1, for some value of 1.

Which means that if we’re serious about X-risk from superviruses, we should be urgently serious about global public health institutions and liability regimes that disincentivize the creation of superviruses more generally. Those rules and institutions should include the actions of LHDEULLMPA, but should not be exclusive to them!

To the extent that there is effective international law (which there isn’t), this is what we do with all sorts of dangerous behaviors. One of the things that makes law ineffective is lack of enforcement, which is certainly an issue because of the lack of international cooperation around anything important. We are slow-walking into World War 3, after all. World War 3 may well be a Hazard with existential implications for the world, but is “AI” the cause of World War 3? I do not think that LHDEULLMPA are the causes of World War 3. The wars are more attributable to stupidity than intelligence, artificial or not.

Fundamentally, the fears of AI X-risk are not, I think, due to specific threats which have little to do with AI per se. Rather, it has to do with the fear that AI will accomplish something (such as supervirus design) that we would otherwise fundamentally not be able to do, or, otherwise, something else I haven’t mentioned yet.

Geoffrey Hinton was on BBC this morning warning about AI X-risk saying, “If you want to know what not being the apex intelligence is like, go ask a chicken.” He was making the case that an AI system would not be used as a tool by humans to destroy other humans, but rather that it would, out of its own survival instinct, dominate humanity. This is where the “similarly bad outcomes” really start to come out of the woodwork. It’s a different Exposure.

So, the question is, to what extent are LHDEULLMPA (as Hazard) going to dominate humanity (the Exposure)? The question is, again, what is the Vulnerablity?

Once again, this seems like a question for a kind of scientist that is not an LLM benchmarker. Humanity is quite volatile, diverse, and resilient. We might look to the history of inter-human domination, including domination using technical means, to see how successful this has been in the past. If we were serious about this kind of risk, we should be investing more in the social and complexity sciences that would test and develop strategies for resilience to this problem. That would necessarily involve looking at, for example, supply chains of energy and components for LHDEULLMPA and their vulnerability to sabotage by humans. It would also require a hard look at the actual costs of operating an LHDEULLMPA and who would be paying those costs as one of them goes about its domineering way.

For the record, I do care about this kind of science, it’s one of the reasons I’ve been working on scikit-agent. Scikit-agent is a toolkit for studying large, complex, multi-agent systems, some of which might by “AI agents”, but others might be nation states, or individuals, or firms, or whatever. The idea is to develop an expressive modeling framework for this kind of problem. The models will support analytic evaluation, as well as simulation-based analysis. There will be algorithms for fitting models to available data. My goal is for this to be a good tool for institution design, including international AI governance design. But I digress.

The deeper frustration here is that while LHDEULLMPA may be the “AI” du jour, LHDEULLMPA are perhaps only one way to spend billions of dollars in unaccountable, unsupervised computing for potentially nefarious ends. I spoke about the problems of law enforcement — how even if we have good laws, under-enforcement makes them ineffective. Just as an example, the European Union, which is very good at passing laws about technology and data, is notoriously uneven in their enforcement of these laws, because the companies hop into the most friendly enforcement jurisdiction, which is Ireland. In both the EU and the US, enforcement of, say, privacy law comes down to very slow and sporadically applied ex post fines which become a “cost of doing business” for the companies that are held liable. That’s not an effective regime.

“AI” regulation has a lot in common with privacy regulation. “AI” regulation could, if we wanted it to be, about corralling the otherwise unaccountable information flows that are sometimes harmful. This is currently an unsolved problem. Or, “AI” regulation could be about guaranteeing that the LHDEULLMPA are aligned with — loyal to — their principals (who are most often not genocidal bioterrorists). There is a legal framework for incentivizing agents to be loyal — fiduciary duties. But interestingly, among those that have written about fiduciary AI, there’s a sizable contingent that are pessimistic about using them as real legal requirements with teeth, in part because of pessimism about the technology industry allowing those regulations to take hold. This is despite the fact that some sectors already have these duties in place, and companies operating in those spaces are working hard for their agents to be compliant!

Which gets us, once again, to the frustrating impasse. The AI X-riskers are telling us that there’s a 10% chance of human extinction because of advanced AI and “want to slow down”. There are policy tools available that would bind the actors that are building and furnishing the LHDEULLMPA that are available on the market. But those same actors are doing their best to crush substantive regulation which would slow down development. Once again, the crux of the problem is not LLM benchmarking, but political economy. We do not (yet) have benchmarks for how well LHDEULLMPA perform on political economy. If we did, we might have a much better sense of how much of a hazard they pose to human domination. On the other hand, it may well be that the political economy of AI regulation is after all partly observable, in so far as the contributions to super PACs that oppose it are public. And it may not, at the end of the day, be a particularly hard computational problem to understand. The roots seem to have already been outlined in social theory written in the 1940s. And yet, for some reason, we keep trying to make this about the LHDEULLMPA.

The Open Silicon Fallacy: How Panic Over Chinese Open-Weight Models Is Distorting American AI Policy

The following article was written with Gemini Flash 3.6 “in the style of The Economist’s Bagehot”. It is an experiment in AI writing. The arguments and structure are mine. – S

I. The Dragons in the Weights

In the summer of 2026, the fashionable anxiety in Washington and Silicon Valley is that Western technological supremacy has been undermined by a collection of downloadable matrix files. When Chinese research labs like DeepSeek, Moonshot AI, and Z.ai released the weights for their latest systems, the initial response from American technology executives was a nervous cough about benchmark scores. When developers discovered these open-weight models could execute multi-step reasoning at a fraction of the cost of renting access to proprietary American cloud endpoints, the nervous cough turned into a geopolitical emergency.

The narrative is simple, compelling, and decidedly panicked. Chinese open-weight architectures are closing the benchmark gap with closed Western labs. Engineering teams weary of driving a metaphorical Ferrari to Whole Foods for basic data transformations are migrating toward open models hosted locally or behind private firewalls.

To visit Capitol Hill today is to encounter three distinct pillars of alarm regarding these releases. First comes the fear of lost technological primacy, as the arrival of competitive Chinese models shatters the illusion that GPU export restrictions would maintain a multi-generational lead. Second comes the specter of standards hegemony, with strategists dreading a world where global software standardizes on Chinese-developed open architectures. Third, and most loudly invoked, comes the proliferation of dual-use capabilities, as downloadable intelligence primitives escape centralized censorship and safety filters.

Yet what is being distributed from Beijing is not open source in the classic sense of transparent codebases and reproducible pipelines. These are open weights—the pre-computed mathematical parameters of deep neural networks whose underlying data and alignment routines remain closely guarded secrets. That they are nevertheless transforming global software architecture says far more about the economics of general computing than about the ideological triumph of open-source idealism.

II. Bootleggers, Baptists, and Moats

It is a sound rule of political economy that whenever a commercial sector and a national security agency begin using identical language to describe a threat, one should look closely at who stands to profit from the proposed remedy.

What Washington is currently witnessing is a classic demonstration of the bootleggers and baptists dynamic. The baptists are the national security hawks, genuinely concerned with technological primacy and cyber-resilience. The bootleggers are the dominant proprietary API vendors, whose staggering valuations depend entirely on convincing the market that intelligence can only be safely consumed as a paid subscription utility.

When proprietary labs lobby for mandatory model licensing, pre-deployment government safety audits, or outright bans on foreign open-weight distributions, they do so under the pious banner of national defense. Yet the operational effect of such proposals is unmistakable, erecting regulatory moats that entrench an oligopolistic duopoly. By framing open-weight distribution as an inherent national security threat, incumbents seek to achieve through regulatory capture what they struggle to maintain through market competition: the enclosure of the AI stack. The public interest is conflated with the profit margins of cloud providers, while the broader software ecosystem is instructed to accept API lock-in as a patriotic duty.

III. The Halloween Documents Revisited

History does not repeat itself, but software executives certainly recycle their memo templates. In the late 1990s, when Microsoft felt its desktop monopoly threatened by Linux, its executives authored internal strategic assessments—the famous Halloween Documents—warning that open software presented a systemic threat to software stability, intellectual property, and commercial viability.

The current rhetoric against open-weight AI models reproduces this playbook line for line. Once again, open distribution is framed as an irresponsible hazard; once again, security through obscurity is held up as the only responsible posture.

Yet incumbent resistance follows a predictable lifecycle, beginning with initial ridicule, progressing to intense alarmism, moving to lobbying for legal restriction, and ultimately settling into a quiet, pragmatic pivot to co-optation. Two decades after penning memos declaring open software an existential cancer, Microsoft spent 7.5 billion dollars to acquire GitHub, transforming itself into the world’s largest host of open-source code. The very proprietary companies that once swore open software was a menace today run their cloud empires on open infrastructure.

IV. How the Defense State Learned to Love the Kernel

The irony of the current policy panic is that the national security establishment has already solved this problem once before. In the early days of networked computing, defense agencies viewed open-source software with profound suspicion, assuming closed, proprietary systems were superior because their source code was hidden behind non-disclosure agreements and commercial firewalls.

By the early 2000s, however, military and intelligence strategists realized that proprietary vendors could not patch vulnerabilities or adapt to new threats as rapidly as a global community of developers inspecting open code. The intelligence community embraced the doctrine of security through visibility. In 2000, the National Security Agency took an open-source Linux kernel, added mandatory access controls directly into its architecture, and handed Security-Enhanced Linux back to the public.

National security interests accommodated open source not by suppressing it, but by co-opting, hardening, and building on top of it. They recognized that controlling an open, auditable standard offered greater agility and defense-in-depth than relying on a commercial black box.

V. Compilers, Not Missiles

The current attempt to govern AI safety by restricting model weights rests on a fundamental misapprehension of what a large language model actually is. Policy makers consistently treat probabilistic language models as if they were self-contained, autonomous products or guided weapons systems that can be aligned at the factory and locked in a box. In reality, a foundation model is a general-purpose computing primitive, serving as the statistical equivalent of a C compiler or an arithmetic logic unit for natural language and code.

Trying to enforce safety at the weight level is as ineffective as trying to secure an operating system by banning specific sequences of assembly language instructions. Raw inference is inherently difficult to control at the parameter level; a model that can write a Python script for a database query can, with minimal prompting, write a script to probe a network port.

Opponents of open weights often argue that publicly downloadable parameters allow offline execution, rendering traditional hardware tracking obsolete. This argument, however, confuses hardware tracking with runtime deployment governance. The true execution boundary is not the local matrix multiplication happening on a graphics card, but the point where an AI system interacts with the real world through API keys, database access, tool-use privileges, and execution environments. Smart policy does not attempt to police raw matrix math in memory; it enforces strict sandboxing, zero-trust permissions, and identity verification at the application layer where actions occur.

VI. Critiquing “Openness”

To defend the availability of open weights is not to be naive about their limitations. Indeed, neural network weights occupy a strange conceptual middle ground. They are not traditional open source code, nor are they merely untrusted software binaries; they are dense, highly compressed mathematical encodings of vast cultural, technical, and linguistic corpora. They cannot be read in any conventional human sense like C instructions, yet they contain whole libraries of distilled human data.

This epistemic opacity is precisely why open weights are necessary. Having direct access to raw model parameters is the strict prerequisite for mechanistic interpretability, local safety probing, and post-hoc red-teaming. A proprietary API offers zero visibility into what lies behind the endpoint, whereas an open weight file allows security researchers to inspect internal activation patterns, trace knowledge representations, and strip out toxic behaviors.

Similarly, open-weight models originating from authoritarian states are indisputably shaped by domestic censorship and potential state alignment. The correct operational response, however, is to treat them with the same caution accorded to foreign-built infrastructure. Western developers can download, audit, strip out state-imposed guardrails, and repurpose foreign base parameters for independent domestic use, turning foreign releases to Western defensive advantage.

As scholars David Gray Widder, Meredith Whittaker, and Sarah Myers West point out in Nature (2024), tech giants frequently engage in openwashing—releasing model weights as a public relations gesture while keeping training datasets, filtering pipelines, and compute infrastructure firmly closed. This critique is vital, but its policy conclusion must be drawn carefully. That open-weight releases represent an incomplete form of openness is an argument for demanding greater transparency and public investment in shared compute and datasets. It is emphatically not an argument for retreating into the arms of proprietary API monopolies.

VII. The Open Security Imperative

The present impulse to restrict, license, or ban open-weight AI models repeats the classical errors of past technological panics. It mistakes corporate rent-seeking for national defense, confuses general computing primitives with finished weapons, and trades long-term systemic resilience for the illusion of central control.

A pragmatic blueprint for AI policy must start from a posture of realism. General-purpose reasoning parameters will circulate globally across open networks regardless of administrative bans. The defense of critical infrastructure relies on open access to model parameters, enabling global researchers to discover vulnerabilities and build defensive countermeasures faster than adversaries can exploit them. Policy must focus its regulatory instruments on the environment where software acts—governing identity verification, agentic tool permissions, data access, and sandboxed execution layers—rather than attempting to criminalize the distribution of general-purpose math.

To lock down American AI within proprietary walled gardens out of fear of foreign open-weight competition would be a historic miscalculation. In the long struggle for technological adaptability and national security, open systems remain, as they have always been, the ultimate line of defense.

References

LLMs as computation

LLMs are now”doing” a lot of technical system design and are the object of a great deal of computer science research. However, I’ve surprised by much of the research that crosses my way (admittedly likely not a great sample) treats LLMs as a general form of intelligence without treating it as a form of computation. I expect that some combination of theory of computation (such as algorithmic information theory) and structural economics is needed to get a rigorous handle on the AI economy. This blog post contains some notes toward this end.

As we all know, an LLM is a collection of neural network weights, trained on a massive amount of information, which consumes tokens and emits predicted next tokens. Simplifying a bit, we can model an LLM as a machine that, given a string of tokens, emits a string of tokens.

Let Σ\Sigma be the set of tokens, Σ\Sigma^* be the space of token strings of any length. Perhaps an LLM is a function:

L:ΣΣL: \Sigma^* \rightarrow \Sigma^*

Really, this is LLM “inference”. I’m omitting the inherent stochasticity of LLMs — more realistically, LL would be a conditional probability distribution. But leave that aside for now.

Assuming that LL can consume as input any string, and in principle produce as output any string, what we have here is a class of “universal programming language”, another formal mathematical construct. “universal programming languages” appear in algorithmic information theory.

The simplest form of “universal programming language” is the print function. It repeats as output anything put into it. People (including myself) once joked that LLMs are glorified autocomplete; they clearly do more than this. The weights must matter.

Really, LLMs are parameterized functions — the parameters θ\theta are weights of the neural network.

Lθ:ΣΣL_\theta: \Sigma^* \rightarrow \Sigma^*

The weights are a compression of a great deal of training data 𝐃\mathbf{D}. Let’s assume training has converted this data to a set of weights T(𝐃)𝛉T(\mathbf{D}) \rightarrow \mathbf{\theta}. We can refer to this foundation model as 𝐋𝛉\mathbf{L_\theta} or 𝐋𝐃\mathbf{L_D}.

What else can you do with these models? You can provide them ‘context’ — additional strings as input. You can fine-tune them on more data. And you can use them for ‘reasoning’ by chaining inputs and outputs.

  • Context: allow multiple string inputs Lθ(c,i)oL_\theta(c, i) \rightarrow o
  • Fine-tuning: T(LD,d)LD+dT(L_D, d) \rightarrow L_{D + d} — further compresses additional data dd into the model weights
  • Reasoning: Lθn(i)Lθ(Lθ(...(Lθ(i)))oL^n_\theta(i) \rightarrow L_\theta(L_\theta(…(L_\theta(i))) \rightarrow o applies the model recursively nn times

So if we want to look at the data and computation pipeline of an LLM based system, we get something like:

(Tn(D,d1,...dn))m(c,i)o(T^n(D,d_1, …d_n))^m(c,i) \rightarrow o

I.e., we train on a base data set and several fine-tuning data sets, pick context and an input, and run inference some number of times. Each of these steps has a cost function, and we can then computer the average costs of solving various sets of problems given the available data, and other statistics. This then can be used to design the most efficient pipelines and markets.

I would be interested in hearing from anybody about whether and how this faithfully captures the essentials of LLMs as a form of computation. This is my ‘mental model’. I have left out tool use and interactivity, among other things, but those can be added in easily.

Why am I writing this? Because I think that clearly articulating the formal properties of LLMs brings a number of issues to light.

First, it foregrounds the importance of training data. Famously, the transformer architecture is very general, and early innovation in LLMs was largely about scaling it up to greater amounts of data. If we are interested in the behavior of LLMs, the training data and the training algorithm are the parts that are not “black boxes” to the model creators.

As we look at the future of LLMs in the economy, we will be looking at the results of differential access to data, as well as what data is commonly available. This shares a lot of patterns with previous iterations of concerns over “big data”, but this is obscured today because of the charisma of the models themselves.

Second, it makes explicit how information can flow and transform into a system output. The information comes first from training and fine-tuning data, then from context, then from system input. If the training and inference algorithms are general enough, none of the information relevant to a specific task comes from those parts of the system. Those algorithms are ‘general computing’.

Third, it breaks up training and inference. While training and inference are not so different in terms of information flow, they are in practice quite different because of their physical and economic costs. Currently, training is more expensive than inference. So, we see a race to, expensively, train general models with more and more data, so that less and less data is needed in context at inference time, and fewer steps are needed during reasoning. A structural model that distinguishes these can discriminate between several investment hypotheses in this space.

Fourth, by revealing LLMs as a form of general data processing and computation, it deflates (in what I think is a good and necessary way) the tendency to see ‘model evaluations’ as the best way to enforce AI accuracy, fairness, privacy, and so on. My general frustration with the model evaluation literature is that LLMs are that if they are a flavor of universal programming language by design, then there will, by definition, always be a jailbreak or a hallucination available to them. A lot of work on ‘guardrails’ at the model level seems to be about making certain kinds of outputs more difficult or expensive to get. As we’ve seen, there will be open models, and they will get fine-tuned by hobbyists and others to get around the guardrails, and so that’s not going to be an effective strategy long term.

This means that a lot of AI product design and regulation seems to be about shifting around the cost functions for achieving certain kinds of outputs with certain data. If ‘bad’ behaviors are expensive, and ‘good’ behaviors are cheap, then we have, in a sense, succeeded. But this means that the underlying economics must be part of the analysis for it to have forward-going relevance and replicability. Today’s model capabilities are a function of whatever the latest investment — at the training and inference level, as well as the data flow of context and inputs, which may go back into training — is. The entire pipeline produces ‘the intelligence’, and it does so at physical and economic cost. Computer science research, per se, with its focus on the currently available digital artifacts, is not going to achieve lasting results unless it expands its purview to these broader systems and considerations. Likewise, evaluations of models alone will not provide us the reliable theoretical knowledge needed to steer public policy. We must take into account production costs and data pipelines.

updates and stubbornness about superintelligence

We seem to be in a new moment of media excitement about the implications of artificial intelligence. This time, the moment is driven by the experience of software engineers and other knowledge workers who are automating their work with ‘agents’. Clause Code etc. The latest generation of models and services is really good at doing things.

Does this change anything about my “position on AI” and superintelligence in particular?

I wrote a brief paper in 2017 about Bostrom’s Superintelligence argument. I concluded that algorithmic self-improvement at the software level would not produce superintelligence. Rather, intelligence group is limited by data and hardware.

In 2025, this conclusion still holds up, as we’ve seen that the recent impressive advances in AI has depended on tremendous capital expenditure on data centers, high-performing chips, and energy. It also depends on well-publicized efforts to collect all the text known to humankind for training data.

About 8 years ago when I was thinking about this, I wrote a bit about the connection between the Superintelligence argument and the Frankfurt School’s views on instrumental reason and capitalism. The alignment of AI with capital has born out, and has been written about by many others. What is striking about the current moment is just how on-the-nose that alignment is in the US, in terms of the full stack of energy, hardware, models, applications, and then some.

So, so far, no update.

In 2021 I published an article saying that we already had artificial systems with the capacity to outperform individual humans at many tasks. They were and still are called corporations or firms. We also had replaced markets with platform, which are similarly more performant in terms of reducing transaction costs. In that article, Jake Goldenfein and I argue that what ultimately matters are the purposes of the social system that operates the AI technology.

I believe this argument also continues to hold up. The successful models and service we are seeing are corporate accomplishments. The corporation is still the relevant unit of analysis when considering AI.

There are a number of interesting things happening now which I think are undertheorized:

  • What is the real economics of AI, given that the supply chains are so long and complex, consistent of both material and intellectual inputs, and the market for demand is uncertain? This is the trillion dollar question in terms of valuations, and it’s unanswered. The empirics here are not very good because things are far out of equlibrium.
  • Put another way: what does AI mean for the relationships between capital, corporations, labor, and consumers? Some of these relationships are mediated by rules about corporate law, intellectual property and data use, and so are determinable by law rather than technology. Information law therefore is a key point of political intervention in an economic system that is otherwise determined by laws of nature (energy, computation, etc.?

To put it another way: superintelligence has been happening and continues to happen. Some of this is due to laws of nature. But there is still a meaningful point of human intervention, which is the laws of humanity. Designing and implementing those laws well remains an important challenge.

One last thought. I’ve been inspired by Beninger’s The Control Revolution (1986) which is a historical account of the information economy in terms of cybernetics and information theory. You can ask an AI to tell you more about it, but one item comes to mind: that each new information technology first seems to threaten the jobs of people doing information work, and then leads to an expanded number of information jobs. This has to do with the way complexity is and is not managed by the technology. There’s an open question whether this generation of AI is any different. The question is truly open, but my hunch at the moment is that today’s AI systems are creating a lot more complexity than they are controlling. We will see.

What about the loyalty of AI agents in government function?

For the private sector, there is a well-developed theory of loyalty that shows up in fiduciary duties. I and others have argued that these duties of loyalty are the right tool to bring artificial intelligence systems (and the companies that produce them) into “alignment” with those that employ them.

At the time of this writing (the early days of the second Trump administration), there appears to be a movement within the federal government to replace many human bureaucratic functions with AI.

Oddly, it doesn’t seem that government workers have as well-defined (or well-enforced) a sense of who or what they are loyal to as fiduciaries in the private sector. This makes aligning these AI systems even more challenging.

Whereas a doctor or lawyer or trustee can be expected to act in the best interest of their principal, a government worker might be loyal to the state broadly speaking, or to their party, or to their boss, or to their own career. These nuances have long been chalked up to “politics”. But the fissures in the U.S. federal government currently are largely about these divisions of loyalty, and the controversies are largely about conflicts of interest that would be forbidden in a private fiduciary context.

So, while we might wish that a democratically elected government have an affirmative obligation to loyalty and care towards, perhaps, the electorate, with subsidiary duties of confidentiality, disclosure, and so on, that is not legally the case. Instead, there’s a much more complex patchwork of duties and powers, and a set of checks and balances which is increasingly resembling the anarchic state of international relations.

The “oath of office” is perhaps an expression or commitment of loyalty that federal government workers could be held to. As far as I know, this is never directly legally enforced.

Between this inherent ambiguity of loyalty, and further complications brought on by the fact that government AI will in most cases be produced by third parties and procured, not hired, make the AI alignment especially fraught.

Political theories and AI

Through a few new emerging projects and opportunities, I’ve had reason to circle back to the topic of Artificial Intelligence and ethics. I wanted to jot down a few notes as some recent reading and conversations have been clarifying some ideas here.

In my work with Jake Goldenfein on this topic (published 2021), we framed the ethical problem of AI in terms of its challenge to liberalism, which we characterize in terms of individual rights (namely, property and privacy rights), a theory of why the free public market makes the guarantees of these rights sufficient for many social goods, and a more recent progressive or egalitarian tendency. We then discuss how AI technologies challenge liberalism and require us to think about post-liberal configurations of society and computation.

A natural reaction to this paper, especially given the political climate in the United States, is “aren’t the alternatives to liberalism even worse?” and it’s true that we do not in that paper outline an alternative to liberalism which a world with AI might aspire to.

John Mearsheimer’s The Great Delusion: Liberal Dreams and International Realities (2018) is a clearly written treatise on political theory. Mearsheimer rose to infamy in 2022 after the Russian invasion of Ukraine because of widely circulated videos of a lecture in 2015 in which he argued that the fault for Russia’s invasion of Crimea in 2014 was due to U.S. foreign policy. It is because of that infamy that I’ve decided to read The Great Delusion, which was a Financial Times Best Book of 2018. The Financial Times editorials have since turned on Mearsheimer. We’ll see what they say about him in another four years. However politically unpopular he may be, I found his points interesting and have decided to look at his more scholarly work. I have not been disappointed, and find that he clearly articulates political philosophy I will use these articulations. I won’t analyze his international relations theory here.

Putting Mearsheimer’s international relations theories entirely aside for now, I’ve been pleased to find The Great Delusion to be a thorough treatise on political theory, and it goes to lengths in Chapter 3 to describe liberalism as a political theory (which will be its target). Mearsheimer distinguished between four different political ideologies, citing many of their key intellectual proponents.

  • Modus vivendi liberalism. (Locke, Smith, Hayek) A theory committed to individual negative rights, such as private property and privacy, against the impositions by the state. The state should be minimal, a “night watchman”. This can involve skepticism about the ability of reason to achieve consensus about the nature of the good life; political toleration of differences is implied by the guarantee of negative rights.
  • Progressive liberalism. (Rawls) A theory committed to individual rights, including both negative rights and positive rights, which can be in tension. An example positive right is equal opportunity, which requires interventions by the state in order to guarantee. So the state must play a stronger role. Progressive liberalism involves more faith in reason to achieve consensus about the good life, as progressivism is a positive moral view imposed on others.
  • Utilitarianism. (Bentham, Mill) A theory committed to the greatest happiness for the greatest number. Not committed to individual rights, and therefore not a liberalism per se. Utilitarian analysis can argue for tradeoffs of rights to achieve greater happiness, and is collectivist, not in individualist, in the sense that it is concerned with utility in aggregate.
  • Liberal idealism. (Hobson, Dewey) A theory committed to the realization of an ideal society as an organic unity of functioning subsystem. Not committed to individual rights primarily, so not a liberalism, though individual rights can be justified on ideal grounds. Influenced by Hegelian views about the unity of the state. Sometimes connected to a positive view of nationalism.

This is a highly useful breakdown of ideas, which we can bring back to discussions of AI ethics.

Jake Goldenfein and I wrote about ‘liberalism’ in a way that, I’m glad to say, is consistent with Mearsheimer. We too identity right- and left- wing strands of liberalism. I believe our argument about AI’s challenge to liberal assumptions still holds water.

Utilitarianism is the foundation of one of the most prominent versions of AI ethics today: Effective Altruism. Much has been written about Effective Altruism and its relationship to AI Safety research. I have expressed some thoughts. Suffice it to say here that there is a utilitarian argument that ‘ethics’ should be about prioritizing the prevention of existential risk to humanity, because existential catastrophe would prevent the high-utility outcome of humanity-as-joyous-galaxy-colonizers. AI is seen, for various reasons, to be a potential source of catastrophic risk, and so AI ethics is about preventing these outcomes. Not everybody agrees with this view.

For now, it’s worth mentioning that there is a connection between liberalism and utilitarianism through theories of economics. While some liberals are committed to individual rights for their own sake, or because of negative views about the possibility of rational agreement about more positive political claims, others have argued that negative rights and lack of government intervention lead to better collective outcomes. Neoclassical economics has produced theories and ‘proofs’ to this effect, which rely on mathematical utility theory, which is a successor to philosophical utilitarianism in some respects.

It is also the case that a great deal of AI technology and technical practice is oriented around the vaguely utilitarian goals of ‘utility maximization’, though this is more about the mathematical operationalization of instrumental reason and less about a social commitment to utility as a political goal. AI practice and neoclassical economics are quite aligned in this way. If I were to put the point precisely, I’d say that the reality of AI, by exposing bounded rationality and its role in society, shows that arguments that negative rights are sufficient for utility-maximizing outcomes are naive, and so are a disappointment for liberals.

I was pleased that Mearsheimer brought up what he calls ‘liberal idealism’ in his book, despite it being perhaps a digression from his broader points. I have wondered how to place my own work, which draws heavily on Helen Nissenbaum’s theory of Contextual Integrity (CI), which is heavily influenced by the work of Michael Walzer. CI is based on a view of a society composed of separable spheres, which distinct functions and internally meaningful social goods, which should not be directly exchanged or compared. Walzer has been called a communitarian. I suggest that CI might be best seen as a variation of liberal idealism, in that it orients ethics towards a view of society as an idealized organic unity.

If the present reality of AI is so disappointing, then we must try to imagine a better ideal, and work our way towards it. I’ve found myself reading more and more work, such as by Felix Adler and Alain Badiou, that advocate for the need for an ideal model of society. What we currently are missing is a good computational model of such a society which could do for idealism what neoclassical economics did for liberalism. Which is, namely, to create a blueprint for a policy and science of its realization. If we were to apply AI to the problem of ethics, it would be good to use it this way.

Practical social forecasting

I was once long ago asked to write a review of Philip Tetlock’s Expert Political Judgment: How Good Is It? How Can We Know? (2006) and was, like a lot of people, very impressed. If you’re not familiar with the book, the gist is that Tetlock, a psychologist, runs a 20 year study asking everybody who could plausibly be called a “political expert” to predict future events, and then scores them using a very reasonable Bayesian scoring system. He then searches the data for insights about what makes for good political forecasting ability. He finds it to be quite rare, but correlated with humbler and more flexible styles of thinking. Tetlock has gone on to pursue and publish about this line of research. There are now forecasting competitions, and the book Superforecasting. Tetlock has a following.

What I caught my attention in the original book, which was somewhat downplayed in the research program as a whole, is that rather simple statistical models, with two or three regressed variables, performed very well in comparison to even the best human experts. In a Bayesian sense, they were at least as good as the best people. These simple models tended towards guessing something close to the base rate of an event, whereas even the best humans tended to believe their own case-specific reasoning somewhat more than they perhaps should have.

This could be seen as a manifestation of the “bias/variance tradeoff” in (machine and other) learning. A learning system must either have a lot of concentration in the probability mass of its prior (bias) or it must spread this mass quite thin (variance). Roughly, a learning system is a good one for its context if, and maybe only if, its prior is a good enough fit for the environment that it’s in. There’s no free lunch. So the only way to improve social scientific forecasting is to encode more domain specific knowledge into the learning system. Or so I thought until recently.

For the past few years I have been working on computational economics tools that enable modelers to imagine and test theories about the dynamics behind our economic observations. This is a rather challenging and rewarding field to work in, especially right now, when the field of Economics is rapidly absorbing new idea from computer science and statistics. Last August, I had the privilege to attend a summer school and conference on the theme of “Deep Learning for Solving and Estimating Dynamic Models” put on by the Econometric Society DSE Summer School. It was awesome.

The biggest, least subtle, takeaway from the summer school and conference is that deep learning is going to be a big deal for Economics, because these techniques make it feasible to solve and estimate models with much higher dimensionality than has been possible with prior methods. By “solve”, I mean coming to conclusions, for a given model of a bunch of agents interacting with each other through, for example, a market, with some notion of their own reward structure, what the equilibrium dynamics of that system are. Solving these kinds of stochastic dynamic control problems, especially when there is nontrivial endogenous aggregation of agent behavior, is computationally quite difficult. But there are cool ways of encoding the equilibrium conditions of the model, or the optimality conditions of the agents involved, into the loss function of a neural network so that the deep learning training architecture works as a model solver. By “estimate”, I mean identify, for a give model, the parameterization of the model that produces results that make some empirical calibration targets maximally likely.

But maybe more foundationally exciting than seeing these results — which were very great — was the work that demonstrated some practical consequences of the double descent phenomenon in deep learning.

Double descent has been discussed, I guess, since 2018 but it has only recently gotten on my radar. It explains a lot about how and why deep learning has blown so many prior machine learning results out of the water. The core idea is that when a neural network is overparameterized — has so many degrees of freedom that, when trained, it can entirely interpolate (reproduce) the training data — it begins to perform better than any underparameterized model.

The underlying reasons for this are deep and somewhat mysterious. I have an intuition about it that I’m not sure checks out properly mathematically, but I will jot it down here anyway. There are some results suggesting that an infinitely parameterized neural network, of a certain kind, is equivalent to a Gaussian Process, a collection of random variables such that any infinite collection of them is a multivariate normal distribution. If the best model that we can ever train is an even largely and more complex Gaussain process, then this suggests that the Central Limit Theorem is once again the rule that explains the world as we see it, but in a far more textured and interesting way than is obvious. The problem with the Central Limit Theory and normal distributions is that they are not explainable — the explanation for the phenomenon is always a plethora of tiny factors, none of which are sufficient individually. And yet, because it is a foundational mathematical rule, it is always available as an explanation for any phenomenon we can experience. A perfect null hypothesis. Which turns out to be the best forecasting tool available?

It’s humbling material to work with, in any case.

References

Azinovic, Marlon and Gaegauf, Luca and Scheidegger, Simon, Deep Equilibrium Nets (May 24, 2019). Available at SSRN: https://ssrn.com/abstract=3393482 or http://dx.doi.org/10.2139/ssrn.3393482

Kelly, Bryan T. and Malamud, Semyon and Zhou, Kangying, The Virtue of Complexity in Return Prediction (December 13, 2021). Swiss Finance Institute Research Paper No. 21-90, Journal of Finance, forthcoming, Available at SSRN: https://ssrn.com/abstract=3984925 or http://dx.doi.org/10.2139/ssrn.3984925

Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B. and Sutskever, I., 2021. Deep double descent: Where bigger models and more data hurt. Journal of Statistical Mechanics: Theory and Experiment, 2021(12), p.124003.

Thoughts on Fiduciary AI

“Designing Fiduciary Artificial Intelligence”, by myself and David Shekman, is now on arXiv. We’re very excited to have had it accepted to Equity and Access in Algorithms, Mechanisms, and Optimization (EAMMO) ’23, a conference I’ve heard great things about. I hope the work speaks for itself. But I wanted to think “out loud” a moment about how that paper fits into my broader research arc.

I’ve been working in the technology and AI ethics space for several years, and this project sits at the intersection of what I see as several trends through that space:

  • AI alignment with human values and interests as a way of improving the safety of powerful systems, largely coming out of AI research institutes like UC Berkeley’s CHAI and, increasingly, industry labs like OpenAI and Deepmind.
  • Information fiduciary and data loyalty proposals, coming out of “privacy” scholarship. This originates with Jack Balkin, is best articulated by Richards and Hartzog, and has been intellectually engaged by Lina Khan, Julie Cohen, James Grimmelmann, and others. Its strongest legal manifestation so far is probably the E.U.’s Data Governance Act, which comes into effect this year.
  • Contextual Integrity (CI), the theory of technology ethics as contextually appropriate information flow, originating with Helen Nissenbaum. In CI, norms of information flow are legitimized by a social context’s purpose and the ends of those participating within it.

The key intuition is that these three ideas all converge on the problem of designing a system to function in the best interests of some group of people who are the designated beneficiaries in the operational context. Once this common point is recognized, it’s easy to connect the dots between many lines of literature and identify where the open problems are.

The recurring “hard part” of all this is framing the AI alignment problem clearly in terms of the duties of legally responsible actors, while still acknowledging that complying with those duties will increasingly be a matter of technical design. There is a disciplinary tendency in computer science literature to illuminate ethical concepts and translate these into technical requirements. There’s a bit of a disconnect between this literature and the implications for liability of a company that deploys AI, and for obvious reasons it’s rare for industry actors to make this connection clear, opting instead to publicize their ‘ethics’. Legal scholars, on the other hand, are quick to point out “ethics washing”, but tend to want to define regulations as broadly as possible, in order to cover a wide range of technical specifics. The more extreme critical legal scholars in this space are skeptical of any technical effort to guarantee compliance. But this leaves the technical actors with little breathing room or guidance. So these fields often talk past each other.

Fiduciary duties outside of the technical context are not controversial. They are in many ways the bedrock of our legal and economic system, and this can’t be denied with a straight face by any lawyer, corporate director, or shareholder. There is no hidden political agenda in fiduciary duties per se. So as a way to get everybody on the same page about duties and beneficiaries, I think they work.

What is inherently a political issue is whether and how fiduciary duties should be expanded to cover new categories of data technology and AI. We were deliberately agnostic about this point in our recent paper, because the work of the paper is to connect the legal and technical dots for fiduciary AI more broadly. However, at a time when many actors have been calling for more AI and data protection regulation, fiduciary duties are one important option which directly addresses the spirit of many people’s concerns.

My hope is that future work will elaborate on how AI can comply with fiduciary duties in practice, and in so doing show what the consequences of fiduciary AI policies would be. As far as I know, there is no cost benefit analysis (CBA) yet for the passing of data loyalty regulations. If the costs to industry actors were sufficiently light, and the benefits to the public sufficiently high, it might be a way to settle what is otherwise an alarming policy issue.