Digifesto

Tag: artificial intelligence

The Open Silicon Fallacy: How Panic Over Chinese Open-Weight Models Is Distorting American AI Policy

The following article was written with Gemini Flash 3.6 “in the style of The Economist’s Bagehot”. It is an experiment in AI writing. The arguments and structure are mine. – S

I. The Dragons in the Weights

In the summer of 2026, the fashionable anxiety in Washington and Silicon Valley is that Western technological supremacy has been undermined by a collection of downloadable matrix files. When Chinese research labs like DeepSeek, Moonshot AI, and Z.ai released the weights for their latest systems, the initial response from American technology executives was a nervous cough about benchmark scores. When developers discovered these open-weight models could execute multi-step reasoning at a fraction of the cost of renting access to proprietary American cloud endpoints, the nervous cough turned into a geopolitical emergency.

The narrative is simple, compelling, and decidedly panicked. Chinese open-weight architectures are closing the benchmark gap with closed Western labs. Engineering teams weary of driving a metaphorical Ferrari to Whole Foods for basic data transformations are migrating toward open models hosted locally or behind private firewalls.

To visit Capitol Hill today is to encounter three distinct pillars of alarm regarding these releases. First comes the fear of lost technological primacy, as the arrival of competitive Chinese models shatters the illusion that GPU export restrictions would maintain a multi-generational lead. Second comes the specter of standards hegemony, with strategists dreading a world where global software standardizes on Chinese-developed open architectures. Third, and most loudly invoked, comes the proliferation of dual-use capabilities, as downloadable intelligence primitives escape centralized censorship and safety filters.

Yet what is being distributed from Beijing is not open source in the classic sense of transparent codebases and reproducible pipelines. These are open weights—the pre-computed mathematical parameters of deep neural networks whose underlying data and alignment routines remain closely guarded secrets. That they are nevertheless transforming global software architecture says far more about the economics of general computing than about the ideological triumph of open-source idealism.

II. Bootleggers, Baptists, and Moats

It is a sound rule of political economy that whenever a commercial sector and a national security agency begin using identical language to describe a threat, one should look closely at who stands to profit from the proposed remedy.

What Washington is currently witnessing is a classic demonstration of the bootleggers and baptists dynamic. The baptists are the national security hawks, genuinely concerned with technological primacy and cyber-resilience. The bootleggers are the dominant proprietary API vendors, whose staggering valuations depend entirely on convincing the market that intelligence can only be safely consumed as a paid subscription utility.

When proprietary labs lobby for mandatory model licensing, pre-deployment government safety audits, or outright bans on foreign open-weight distributions, they do so under the pious banner of national defense. Yet the operational effect of such proposals is unmistakable, erecting regulatory moats that entrench an oligopolistic duopoly. By framing open-weight distribution as an inherent national security threat, incumbents seek to achieve through regulatory capture what they struggle to maintain through market competition: the enclosure of the AI stack. The public interest is conflated with the profit margins of cloud providers, while the broader software ecosystem is instructed to accept API lock-in as a patriotic duty.

III. The Halloween Documents Revisited

History does not repeat itself, but software executives certainly recycle their memo templates. In the late 1990s, when Microsoft felt its desktop monopoly threatened by Linux, its executives authored internal strategic assessments—the famous Halloween Documents—warning that open software presented a systemic threat to software stability, intellectual property, and commercial viability.

The current rhetoric against open-weight AI models reproduces this playbook line for line. Once again, open distribution is framed as an irresponsible hazard; once again, security through obscurity is held up as the only responsible posture.

Yet incumbent resistance follows a predictable lifecycle, beginning with initial ridicule, progressing to intense alarmism, moving to lobbying for legal restriction, and ultimately settling into a quiet, pragmatic pivot to co-optation. Two decades after penning memos declaring open software an existential cancer, Microsoft spent 7.5 billion dollars to acquire GitHub, transforming itself into the world’s largest host of open-source code. The very proprietary companies that once swore open software was a menace today run their cloud empires on open infrastructure.

IV. How the Defense State Learned to Love the Kernel

The irony of the current policy panic is that the national security establishment has already solved this problem once before. In the early days of networked computing, defense agencies viewed open-source software with profound suspicion, assuming closed, proprietary systems were superior because their source code was hidden behind non-disclosure agreements and commercial firewalls.

By the early 2000s, however, military and intelligence strategists realized that proprietary vendors could not patch vulnerabilities or adapt to new threats as rapidly as a global community of developers inspecting open code. The intelligence community embraced the doctrine of security through visibility. In 2000, the National Security Agency took an open-source Linux kernel, added mandatory access controls directly into its architecture, and handed Security-Enhanced Linux back to the public.

National security interests accommodated open source not by suppressing it, but by co-opting, hardening, and building on top of it. They recognized that controlling an open, auditable standard offered greater agility and defense-in-depth than relying on a commercial black box.

V. Compilers, Not Missiles

The current attempt to govern AI safety by restricting model weights rests on a fundamental misapprehension of what a large language model actually is. Policy makers consistently treat probabilistic language models as if they were self-contained, autonomous products or guided weapons systems that can be aligned at the factory and locked in a box. In reality, a foundation model is a general-purpose computing primitive, serving as the statistical equivalent of a C compiler or an arithmetic logic unit for natural language and code.

Trying to enforce safety at the weight level is as ineffective as trying to secure an operating system by banning specific sequences of assembly language instructions. Raw inference is inherently difficult to control at the parameter level; a model that can write a Python script for a database query can, with minimal prompting, write a script to probe a network port.

Opponents of open weights often argue that publicly downloadable parameters allow offline execution, rendering traditional hardware tracking obsolete. This argument, however, confuses hardware tracking with runtime deployment governance. The true execution boundary is not the local matrix multiplication happening on a graphics card, but the point where an AI system interacts with the real world through API keys, database access, tool-use privileges, and execution environments. Smart policy does not attempt to police raw matrix math in memory; it enforces strict sandboxing, zero-trust permissions, and identity verification at the application layer where actions occur.

VI. Critiquing “Openness”

To defend the availability of open weights is not to be naive about their limitations. Indeed, neural network weights occupy a strange conceptual middle ground. They are not traditional open source code, nor are they merely untrusted software binaries; they are dense, highly compressed mathematical encodings of vast cultural, technical, and linguistic corpora. They cannot be read in any conventional human sense like C instructions, yet they contain whole libraries of distilled human data.

This epistemic opacity is precisely why open weights are necessary. Having direct access to raw model parameters is the strict prerequisite for mechanistic interpretability, local safety probing, and post-hoc red-teaming. A proprietary API offers zero visibility into what lies behind the endpoint, whereas an open weight file allows security researchers to inspect internal activation patterns, trace knowledge representations, and strip out toxic behaviors.

Similarly, open-weight models originating from authoritarian states are indisputably shaped by domestic censorship and potential state alignment. The correct operational response, however, is to treat them with the same caution accorded to foreign-built infrastructure. Western developers can download, audit, strip out state-imposed guardrails, and repurpose foreign base parameters for independent domestic use, turning foreign releases to Western defensive advantage.

As scholars David Gray Widder, Meredith Whittaker, and Sarah Myers West point out in Nature (2024), tech giants frequently engage in openwashing—releasing model weights as a public relations gesture while keeping training datasets, filtering pipelines, and compute infrastructure firmly closed. This critique is vital, but its policy conclusion must be drawn carefully. That open-weight releases represent an incomplete form of openness is an argument for demanding greater transparency and public investment in shared compute and datasets. It is emphatically not an argument for retreating into the arms of proprietary API monopolies.

VII. The Open Security Imperative

The present impulse to restrict, license, or ban open-weight AI models repeats the classical errors of past technological panics. It mistakes corporate rent-seeking for national defense, confuses general computing primitives with finished weapons, and trades long-term systemic resilience for the illusion of central control.

A pragmatic blueprint for AI policy must start from a posture of realism. General-purpose reasoning parameters will circulate globally across open networks regardless of administrative bans. The defense of critical infrastructure relies on open access to model parameters, enabling global researchers to discover vulnerabilities and build defensive countermeasures faster than adversaries can exploit them. Policy must focus its regulatory instruments on the environment where software acts—governing identity verification, agentic tool permissions, data access, and sandboxed execution layers—rather than attempting to criminalize the distribution of general-purpose math.

To lock down American AI within proprietary walled gardens out of fear of foreign open-weight competition would be a historic miscalculation. In the long struggle for technological adaptability and national security, open systems remain, as they have always been, the ultimate line of defense.

References

LLMs as computation

LLMs are now”doing” a lot of technical system design and are the object of a great deal of computer science research. However, I’ve surprised by much of the research that crosses my way (admittedly likely not a great sample) treats LLMs as a general form of intelligence without treating it as a form of computation. I expect that some combination of theory of computation (such as algorithmic information theory) and structural economics is needed to get a rigorous handle on the AI economy. This blog post contains some notes toward this end.

As we all know, an LLM is a collection of neural network weights, trained on a massive amount of information, which consumes tokens and emits predicted next tokens. Simplifying a bit, we can model an LLM as a machine that, given a string of tokens, emits a string of tokens.

Let Σ\Sigma be the set of tokens, Σ\Sigma^* be the space of token strings of any length. Perhaps an LLM is a function:

L:ΣΣL: \Sigma^* \rightarrow \Sigma^*

Really, this is LLM “inference”. I’m omitting the inherent stochasticity of LLMs — more realistically, LL would be a conditional probability distribution. But leave that aside for now.

Assuming that LL can consume as input any string, and in principle produce as output any string, what we have here is a class of “universal programming language”, another formal mathematical construct. “universal programming languages” appear in algorithmic information theory.

The simplest form of “universal programming language” is the print function. It repeats as output anything put into it. People (including myself) once joked that LLMs are glorified autocomplete; they clearly do more than this. The weights must matter.

Really, LLMs are parameterized functions — the parameters θ\theta are weights of the neural network.

Lθ:ΣΣL_\theta: \Sigma^* \rightarrow \Sigma^*

The weights are a compression of a great deal of training data 𝐃\mathbf{D}. Let’s assume training has converted this data to a set of weights T(𝐃)𝛉T(\mathbf{D}) \rightarrow \mathbf{\theta}. We can refer to this foundation model as 𝐋𝛉\mathbf{L_\theta} or 𝐋𝐃\mathbf{L_D}.

What else can you do with these models? You can provide them ‘context’ — additional strings as input. You can fine-tune them on more data. And you can use them for ‘reasoning’ by chaining inputs and outputs.

  • Context: allow multiple string inputs Lθ(c,i)oL_\theta(c, i) \rightarrow o
  • Fine-tuning: T(LD,d)LD+dT(L_D, d) \rightarrow L_{D + d} — further compresses additional data dd into the model weights
  • Reasoning: Lθn(i)Lθ(Lθ(...(Lθ(i)))oL^n_\theta(i) \rightarrow L_\theta(L_\theta(…(L_\theta(i))) \rightarrow o applies the model recursively nn times

So if we want to look at the data and computation pipeline of an LLM based system, we get something like:

(Tn(D,d1,...dn))m(c,i)o(T^n(D,d_1, …d_n))^m(c,i) \rightarrow o

I.e., we train on a base data set and several fine-tuning data sets, pick context and an input, and run inference some number of times. Each of these steps has a cost function, and we can then computer the average costs of solving various sets of problems given the available data, and other statistics. This then can be used to design the most efficient pipelines and markets.

I would be interested in hearing from anybody about whether and how this faithfully captures the essentials of LLMs as a form of computation. This is my ‘mental model’. I have left out tool use and interactivity, among other things, but those can be added in easily.

Why am I writing this? Because I think that clearly articulating the formal properties of LLMs brings a number of issues to light.

First, it foregrounds the importance of training data. Famously, the transformer architecture is very general, and early innovation in LLMs was largely about scaling it up to greater amounts of data. If we are interested in the behavior of LLMs, the training data and the training algorithm are the parts that are not “black boxes” to the model creators.

As we look at the future of LLMs in the economy, we will be looking at the results of differential access to data, as well as what data is commonly available. This shares a lot of patterns with previous iterations of concerns over “big data”, but this is obscured today because of the charisma of the models themselves.

Second, it makes explicit how information can flow and transform into a system output. The information comes first from training and fine-tuning data, then from context, then from system input. If the training and inference algorithms are general enough, none of the information relevant to a specific task comes from those parts of the system. Those algorithms are ‘general computing’.

Third, it breaks up training and inference. While training and inference are not so different in terms of information flow, they are in practice quite different because of their physical and economic costs. Currently, training is more expensive than inference. So, we see a race to, expensively, train general models with more and more data, so that less and less data is needed in context at inference time, and fewer steps are needed during reasoning. A structural model that distinguishes these can discriminate between several investment hypotheses in this space.

Fourth, by revealing LLMs as a form of general data processing and computation, it deflates (in what I think is a good and necessary way) the tendency to see ‘model evaluations’ as the best way to enforce AI accuracy, fairness, privacy, and so on. My general frustration with the model evaluation literature is that LLMs are that if they are a flavor of universal programming language by design, then there will, by definition, always be a jailbreak or a hallucination available to them. A lot of work on ‘guardrails’ at the model level seems to be about making certain kinds of outputs more difficult or expensive to get. As we’ve seen, there will be open models, and they will get fine-tuned by hobbyists and others to get around the guardrails, and so that’s not going to be an effective strategy long term.

This means that a lot of AI product design and regulation seems to be about shifting around the cost functions for achieving certain kinds of outputs with certain data. If ‘bad’ behaviors are expensive, and ‘good’ behaviors are cheap, then we have, in a sense, succeeded. But this means that the underlying economics must be part of the analysis for it to have forward-going relevance and replicability. Today’s model capabilities are a function of whatever the latest investment — at the training and inference level, as well as the data flow of context and inputs, which may go back into training — is. The entire pipeline produces ‘the intelligence’, and it does so at physical and economic cost. Computer science research, per se, with its focus on the currently available digital artifacts, is not going to achieve lasting results unless it expands its purview to these broader systems and considerations. Likewise, evaluations of models alone will not provide us the reliable theoretical knowledge needed to steer public policy. We must take into account production costs and data pipelines.

updates and stubbornness about superintelligence

We seem to be in a new moment of media excitement about the implications of artificial intelligence. This time, the moment is driven by the experience of software engineers and other knowledge workers who are automating their work with ‘agents’. Clause Code etc. The latest generation of models and services is really good at doing things.

Does this change anything about my “position on AI” and superintelligence in particular?

I wrote a brief paper in 2017 about Bostrom’s Superintelligence argument. I concluded that algorithmic self-improvement at the software level would not produce superintelligence. Rather, intelligence group is limited by data and hardware.

In 2025, this conclusion still holds up, as we’ve seen that the recent impressive advances in AI has depended on tremendous capital expenditure on data centers, high-performing chips, and energy. It also depends on well-publicized efforts to collect all the text known to humankind for training data.

About 8 years ago when I was thinking about this, I wrote a bit about the connection between the Superintelligence argument and the Frankfurt School’s views on instrumental reason and capitalism. The alignment of AI with capital has born out, and has been written about by many others. What is striking about the current moment is just how on-the-nose that alignment is in the US, in terms of the full stack of energy, hardware, models, applications, and then some.

So, so far, no update.

In 2021 I published an article saying that we already had artificial systems with the capacity to outperform individual humans at many tasks. They were and still are called corporations or firms. We also had replaced markets with platform, which are similarly more performant in terms of reducing transaction costs. In that article, Jake Goldenfein and I argue that what ultimately matters are the purposes of the social system that operates the AI technology.

I believe this argument also continues to hold up. The successful models and service we are seeing are corporate accomplishments. The corporation is still the relevant unit of analysis when considering AI.

There are a number of interesting things happening now which I think are undertheorized:

  • What is the real economics of AI, given that the supply chains are so long and complex, consistent of both material and intellectual inputs, and the market for demand is uncertain? This is the trillion dollar question in terms of valuations, and it’s unanswered. The empirics here are not very good because things are far out of equlibrium.
  • Put another way: what does AI mean for the relationships between capital, corporations, labor, and consumers? Some of these relationships are mediated by rules about corporate law, intellectual property and data use, and so are determinable by law rather than technology. Information law therefore is a key point of political intervention in an economic system that is otherwise determined by laws of nature (energy, computation, etc.?

To put it another way: superintelligence has been happening and continues to happen. Some of this is due to laws of nature. But there is still a meaningful point of human intervention, which is the laws of humanity. Designing and implementing those laws well remains an important challenge.

One last thought. I’ve been inspired by Beninger’s The Control Revolution (1986) which is a historical account of the information economy in terms of cybernetics and information theory. You can ask an AI to tell you more about it, but one item comes to mind: that each new information technology first seems to threaten the jobs of people doing information work, and then leads to an expanded number of information jobs. This has to do with the way complexity is and is not managed by the technology. There’s an open question whether this generation of AI is any different. The question is truly open, but my hunch at the moment is that today’s AI systems are creating a lot more complexity than they are controlling. We will see.

I’m building something new

Push has come to shove, and I have started building something new.

I’m not building it alone, thank God, but it’s a very petite open source project at the moment.

I’m convinced that it is actually something new. Some of my colleagues are excited to hear that what I’m building will exist for them shortly. That’s encouraging! I also tried to build this thing in the context of another project that’s doing something similar. I was told that I was being too ambitious and couldn’t pull it off. That wasn’t exactly encouraging, but was good evidence that I’m actually doing something new. I will now do this thing and take the credit. So much the better for me.

What is it that I’m building?

Well, you see, it’s software for modeling economic systems. Or sociotechnical systems. Or, broadly speaking, complex systems with agents in them. Also, doing statistics with those models: fitting them to data, understanding what emergent properties occur in them, exploring counterfactuals, and so on.

I will try to answer some questions I wish somebody would ask me.

Q: Isn’t that just agent-based modeling? Why aren’t you just using NetLogo or something?

A: Agent-based modeling (ABM) is great, but it’s a very expansive term that means a lot of things. Very often, ABMs consist of agents whose behavior is governed by simple rules, rather directed towards accomplishing goals. That notion of “agent” in ABM is almost entirely opposed to the notion of “agent” used in AI — propagated by Stuart Russell, for example. To AI people, goal-directedness is essential for agency. I’m not committed to rational behavior in this framework — I’m not an economist! But I think a requirement to be able to train agents’ decision rules with respect to their goals.

There are a couple other ways in which I’m not doing paradigmatic ABM with this project. One is that I’m not focused on agents moving in 2D or 3D space. Rather, I’m much more interested in the settings defined by systems of structural equations. So, more continuous state spaces. I’m basing this work on years of contributing to heterogeneous agent macroeconomics tooling, and my frustrating with that paradigm. So, no turtles on patches. I anticipate spatial and even geospatial extensions to what I’m building would be really cool and useful. But I’m not there yet.

I think what I’m working on is ABM in the extended sense that Rob Axtell and Doyne Farmer use the term, and I hope to one day show them what I’m doing and for them to think it’s cool.

Q: Wait, is this about AI agents, as in Generative AI?

A: Ahaha… mainly no, but a little yes. I’m talking about “agents” in the more general sense used before the GenAI industry tried to make the word about them. I don’t see Generative AI or LLMs to be a fundamental part of what I’m building. However, I do see what I’m building as a tool for evaluating the economic impact and trustworthiness of GenAI systems by modeling their supply chains and social consequences. And I can imagine deeper integrations with “(generative) agentic AI” down the line. I am building a tool, and an LLM might engage it through “tool use”. It’s also I suppose possible to make the agents inside the model use LLMs somehow, though I don’t see a good reason for that at the moment.

Q: Does it use AI at all?

A: Yes! I mean, perhaps you know that “AI” has meant many things and much of what it has meant is now considered quite mundane. But it does use deep learning, which is something at “AI” means now. In particular, part of the core functionality that I’m trying to build into it is a flexible version of the deep learning econometrics methods invented not-too-long-ago by Lilia and Serguei Maliar. I hope to one day show this project to them, and for them to think it’s cool. Deep learning methods have become quite popular in economics, and this is in some ways yet-another-deep-learning-economics project. I hope it has a few features that distinguish it.

Q: How is this different from somebody else’s deep learning economics analysis package?

A: Great question! There are a few key ways that it’s different. One is that it’s designed around a clean separation between model definition and solution algorithms. There will be no model-specific solution code in this project. It’s truly intended to be library, comparable to scikit-learn, but for systems of agents. In fact, I’m calling this project scikit-agent. You heard it hear first!

Separating the model definitions from the solution algorithms means that there’s a lot more flexibility in how models are composed. This framework is based on the idea that parts of a model can be “blocks” which can be composed into more complex models. The “blocks” are bundles of structural equations, which can include state, control, and reward variables.

These ‘blocks’ are symbolically defined systems or environments. “Solving” the agent strategies in the multi-agent environment will be done with deep learning, otherwise known as artificial neural networks. So I think that it will be fair to call this framework a “neurosymbolic AI system”. I hope that saying that makes it easier to find funding for it down the line :)

Q: That sounds a little like causal game theory, or multi-agent influence diagrams. Are those part of this?

A: In fact, yes, so glad you asked. I think there’s a deep equivalence between multi-agent influence diagrams and ABM/computational economics which hasn’t been well explored. There are small notational differences that keep these communities from communicating. There are also a handful of substantively difficult theoretical issues that need to be settled with respect to, say, under what conditions a dynamic structure causal game can be solved using multi-agent reinforcement learning. These are cool problems, and I hope the thing I’m building implements good solutions to them.

Q: So, this is a framework for modeling dynamic Pearlian causal models, with multiple goal-directed agents, solving those models for agent strategy, and using those model econometrically?

A: Exactly.

Q: Does the thing you are building have any practical value? Or is it just more weird academic code?

A: I have high hopes that this thing I’m building could have a lot of practical value. Rigorous analysis of complex sociotechnical and economic systems remain hard problems. In finance, for example, as well as public policy, insurance, international relations, and other fields. I do hope what I’m building interfaces well with real data to help with significant decision-making. These are problems that Generative AI is quite bad at, I believe. I’m trying to build a strong, useful foundation for working with statistical models that include agents in them. This is more difficult than regression or even ‘transformer’-based learning from media, because the agents are solving optimization problems inside the model.

Q: What are the first applications you have in mind for this tool?

A: I’m motivated to build this because I think it’s needed to address questions in technology policy and design. This is the main subject of my NSF-funded research over the past several years. Here are some problems I’m actively working on which overlap with the scope of this tool:

  • Integrating Differential Privacy and Contextual Integrity. I have a working paper with Rachel Cummings where we use Structural Causal Games (SCGs) to set up the parameter tuning of a differentially private system as a mechanism design problem. The tool I’m building will be great at representing structural causal games (SCG). With it, the theoretical technique can be used in practice.
  • Understanding the Effects of Consumer Finance Policy. I’m working with an amazing team on a project modeling consumer lending and the effects of various consumer protection regulations. We are looking at anti-usury laws, nondiscrimination rules, forgiveness of negative information, the use of alternative data by fintech companies, and so on. This policy analysis involves comparing a number of different scenarios and looking for what regulations produce which results, robustly. I’m building a tool to solve this problem.
  • AI Governance and Fiduciary Duties. A lot of people have pointed out that effective AI governance requires an understanding of AI supply chains. AI services rely on complex data flows through multiple actors, which are often imperfectly aligned and incompletely contracted. They also depend on physical data centers and the consumption of energy. This raises many questions around liability and quality control that wind up ultimately being about institutional design rather than neural network architectures. I’m building a tool to help reason through these institution design questions. In other words, I’m building a new way to do threat modeling on AI supply chain ecosystems.

Q: Really?

A: Yes! I sometimes have a hard time wrapping my own head around what I’m doing, which is why I’ve written out this blog post. But I do feel very good about what I’m working on at the moment. I think it has a lot of potential.

Notes about “Data Science and the Decline of Liberal Law and Ethics”

Jake Goldenfein and I have put up on SSRN our paper, “Data Science and the Decline of Liberal Law and Ethics”. I’ve mentioned it on this blog before as something I’m excited about. It’s also been several months since we’ve finalized it, and I wanted to quickly jot some notes about it based on considerations going into it and since then.

The paper was the result of a long and engaged collaboration with Jake which started from a somewhat different place. We considered the question, “What is sociopolitical emancipation in the paradigm of control?” That was a mouthful, but it captured what we were going for:

  • Like a lot of people today, we are interested in the political project of freedom. Not just freedom in narrow, libertarian senses that have proven to be self-defeating, but in broader senses of removing social barriers and systems of oppression. We were ambivalent about the form that would take, but figured it was a positive project almost anybody would be on board with. We called this project emancipation.
  • Unlike a certain prominent brand of critique, we did not begin from an anthropological rejection of the realism of foundational mathematical theory from STEM and its application to human behavior. In this paper, we did not make the common move of suggesting that the source of our ethical problems is one that can be solved by insisting on the terminology or methodological assumptions of some other discipline. Rather, we took advances in, e.g., AI as real scientific accomplishments that are telling us how the world works. We called this scientific view of the world the paradigm of control, due to its roots in cybernetics.

I believe our work is making a significant contribution to the “ethics of data science” debate because it is quite rare to encounter work that is engaged with both project. It’s common to see STEM work with no serious moral commitments or valence. And it’s common to see the delegation of what we would call emancipatory work to anthropological and humanistic disciplines: the STS folks, the media studies people, even critical X (race, gender, etc.) studies. I’ve discussed the limitations of this approach, however well-intentioned, elsewhere. Often, these disciplines argue that the “unethical” aspect of STEM is because of their methods, discourses, etc. To analyze things in terms of their technical and economic properties is to lose the essence of ethics, which is aligned with anthropological methods that are grounded in respectful, phenomenological engagement with their subjects.

This division of labor between STEM and anthropology has, in my view (I won’t speak for Jake) made it impossible to discuss ethical problems that fit uneasily in either field. We tried to get at these. The ethical problem is instrumentality run amok because of the runaway economic incentives of private firms combined with their expanded cognitive powers as firms, a la Herbert Simon.

This is not a terribly original point and we hope it is not, ultimately, a fringe political position either. If Martin Wolf can write for the Financial Times that there is something threatening to democracy about “the shift towards the maximisation of shareholder value as the sole goal of companies and the associated tendency to reward management by reference to the price of stocks,” so can we, and without fear that we will be targeted in the next red scare.

So what we are trying to add is this: there is a cognitivist explanation for why firms can become so enormously powerful relative to individual “natural persons”, one that is entirely consistent with the STEM foundations that have become dominant in places like, most notably, UC Berkeley (for example) as “data science”. And, we want to point out, the consequences of that knowledge, which we take to be scientific, runs counter to the liberal paradigm of law and ethics. This paradigm, grounded in individual autonomy and privacy, is largely the paradigm animating anthropological ethics! So we are, a bit obliquely, explaining why the the data science ethics discourse has gelled in the ways that it has.

We are not satisfied with the current state of ‘data science ethics’ because to the extent that they cling to liberalism, we fear that they miss and even obscure the point, which can best be understood in a different paradigm.

We left as unfinished the hard work of figuring out what the new, alternative ethical paradigm that took cognitivism, statistics, and so on seriously would look like. There are many reasons beyond the conference publication page limit why we were unable to complete the project. The first of these is that, as I’ve been saying, it’s terribly hard to convince anybody that this is a project worth working on in the first place. Why? My view of this may be too cynical, but my explanations are that either (a) this is an interdisciplinary third rail because it upsets the balance of power between different academic departments, or (b) this is an ideological third rail because it successfully identifies a contradiction in the current sociotechnical order in a way that no individual is incentivized to recognize, because that order incentivizes individuals to disperse criticism of its core institutional logic of corporate agency, or (c) it is so hard for any individual to conceive of corporate cognition because of how it exceeds the capacity of human understanding that speaking in this way sounds utterly speculative to a lot of fo people. The problem is that it requires attributing cognitive and adaptive powers to social forms, and a successful science of social forms is, at best, in the somewhat gnostic domain of complex systems research.

The latter are rarely engaged in technology policy but I think it’s the frontier.

References

Benthall, Sebastian and Goldenfein, Jake, Data Science and the Decline of Liberal Law and Ethics (June 22, 2020). Ethics of Data Science Conference – Sydney 2020 (forthcoming). Available at SSRN: https://ssrn.com/abstract=

computational institutions as non-narrative collective action

Nils Gilman recently pointed to a book chapter that confirms the need for “official futures” in capitalist institutions.

Nils indulged me in a brief exchange that helped me better grasp at a bothersome puzzle.

There is a certain class of intellectuals that insist on the primacy of narratives as a mode of human experience. These tend to be, not too surprisingly, writers and other forms of storytellers.

There is a different class of intellectuals that insists on the primacy of statistics. Statistics does not make it easy to tell stories because it is largely about the complexity of hypotheses and our lack of confidence in them.

The narrative/statistic divide could be seen as a divide between academic disciplines. It has often been taken to be, I believe wrongly, the crux of the “technology ethics” debate.

I questioned Nils as to whether his generalization stood up to statistically driven allocation of resources; i.e., those decisions made explicitly on probabilistic judgments. He argued that in the end, management and collective action require consensus around narrative.

In other words, what keeps narratives at the center of human activity is that (a) humans are in the loop, and (b) humans are collectively in the loop.

The idea that communication is necessary for collective action is one I used to put great stock in when studying Habermas. For Habermas, consensus, and especially linguistic consensus, is how humanity moves together. Habermas contrasted this mode of knowledge aimed at consensus and collective action with technical knowledge, which is aimed at efficiency. Habermas envisioned a society ruled by communicative rationality, deliberative democracy; following this line of reasoning, this communicative rationality would need to be a narrative rationality. Even if this rationality is not universal, it might, in Habermas’s later conception of governance, be shared by a responsible elite. Lawyers and a judiciary, for example.

The puzzle that recurs again and again in my work has been the challenge of communicating how technology has become an alternative form of collective action. The claim made by some that technologists are a social “other” makes more sense if one sees them (us) as organizing around non-narrative principles of collective behavior.

It is I believe beyond serious dispute that well-constructed, statistically based collective decision-making processes perform better than many alternatives. In the field of future predictions, Phillip Tetlock’s work on superforecasting teams and prior work on expert political judgment has long stood as an empirical challenge to the supposed primacy of narrative-based forecasting. This challenge has not been taken up; it seems rather one-sided. One reason for this may be because the rationale for the effectiveness of these techniques rests ultimately in the science of statistics.

It is now common to insist that Artificial Intelligence should be seen as a sociotechnical system and not as a technological artifact. I wholeheartedly agree with this position. However, it is sometimes implied that to understand AI as a social+ system, one must understand it one narrative terms. This is an error; it would imply that the collective actions made to build an AI system and the technology itself are held together by narrative communication.

But if the whole purpose of building an AI system is to collectively act in a way that is more effective because of its facility with the nuances of probability, then the narrative lens will miss the point. The promise and threat of AI is that is delivers a different, often more effective form of collective or institution. I’ve suggested that computational institution might be the best way to refer to such a thing.

The Data Processing Inequality and bounded rationality

I have long harbored the hunch that information theory, in the classic Shannon sense, and social theory are deeply linked. It has proven to be very difficult to find an audience for this point of view or an opportunity to work on it seriously. Shannon’s information theory is widely respected in engineering disciplines; many social theorists who are unfamiliar with it are loathe to admit that something from engineering should carry essential insights for their own field. Meanwhile, engineers are rarely interested in modeling social systems.

I’ve recently discovered an opportunity to work on this problem through my dissertation work, which is about privacy engineering. Privacy is a subtle social concept but also one that has been rigorously formalized. I’m working on formal privacy theory now and have been reminded of a theorem from information theory: the Data Processing Theorem. What strikes me about this theorem is that is captures an point that comes up again and again in social and political problems, though it’s a point that’s almost never addressed head on.

The Data Processing Inequality (DPI) states that for three random variables, X, Y, and Z, arranged in Markov Chain such that X \rightarrow Y \rightarrow Z, then I(X,Z) \leq I(X,Y), where here I stands for mutual information. Mutual information is a measure of how much two random variables carry information about each other. If $I(X,Y) = 0$, that means the variables are independent. $I(X,Y) \geq 0$ always–that’s just a mathematical fact about how it’s defined.

The implications of this for psychology, social theory, and artificial intelligence are I think rather profound. It provides a way of thinking about bounded rationality in a simple and generalizable way–something I’ve been struggling to figure out for a long time.

Suppose that there’s a big world out the, W and there’s am organism, or a person, or a sociotechnical organization within it, Y. The world is big and complex, which implies that it has a lot of informational entropy, H(W). Through whatever sensory apparatus is available to Y, it acquires some kind of internal sensory state. Because this organism is much small than the world, its entropy is much lower. There are many fewer possible states that the organism can be in, relative to the number of states of the world. H(W) >> H(Y). This in turn bounds the mutual information between the organism and the world: I(W,Y) \leq H(Y)

Now let’s suppose the actions that the organism takes, Z depend only on its internal state. It is an agent, reacting to its environment. Well whatever these actions are, they can only be so calibrated to the world as the agent had capacity to absorb the world’s information. I.e., I(W,Z) \leq H(Y) << H(W). The implication is that the more limited the mental capacity of the organism, the more its actions will be approximately independent of the state of the world that precedes it.

There are a lot of interesting implications of this for social theory. Here are a few cases that come to mind.

I've written quite a bit here (blog links) and here (arXiv) about Bostrom’s superintelligence argument and why I’m generally not concerned with the prospect of an artificial intelligence taking over the world. My argument is that there are limits to how much an algorithm can improve itself, and these limits put a stop to exponential intelligence explosions. I’ve been criticized on the grounds that I don’t specify what the limits are, and that if the limits are high enough then maybe relative superintelligence is possible. The Data Processing Inequality gives us another tool for estimating the bounds of an intelligence based on the range of physical states it can possibly be in. How calibrated can a hegemonic agent be to the complexity of the world? It depends on the capacity of that agent to absorb information about the world; that can be measured in information entropy.

A related case is a rendering of Scott’s Seeing Like a State arguments. Why is it that “high modernist” governments failed to successfully control society through scientific intervention? One reason is that the complexity of the system they were trying to manage vastly outsized the complexity of the centralized control mechanisms. Centralized control was very blunt, causing many social problems. Arguably, behavioral targeting and big data centers today equip controlling organizations with more informational capacity (more entropy), but they
still get it wrong sometimes, causing privacy violations, because they can’t model the entirety of the messy world we’re in.

The Data Processing Inequality is also helpful for explaining why the world is so messy. There are a lot of different agents in the world, and each one only has so much bandwidth for taking in information. This means that most agents are acting almost independently from each other. The guiding principle of society isn’t signal, it’s noise. That explains why there are so many disorganized heavy tail distributions in social phenomena.

Importantly, if we let the world at any time slice be informed by the actions of many agents acting nearly independently from each other in the slice before, then that increases the entropy of the world. This increases the challenge for any particular agent to develop an effective controlling strategy. For this reason, we would expect the world to get more out of control the more intelligence agents are on average. The popularity of the personal computer perhaps introduced a lot more entropy into the world, distributed in an agent-by-agent way. Moreover, powerful controlling data centers may increase the world’s entropy, rather than redtucing it. So even if, for example, Amazon were to try to take over the world, the existence of Baidu would be a major obstacle to its plans.

There are a lot of assumptions built into these informal arguments and I’m not wedded to any of them. But my point here is that information theory provides useful tools for thinking about agents in a complex world. There’s potential for using it for modeling sociotechnical systems and their limitations.

Managerialism as political philosophy

Technologically mediated spaces and organizations are frequently described by their proponents as alternatives to the state. From David Clark’s maxim of Internet architecture, “We reject: kings, presidents and voting. We believe in: rough consensus and running code”, to cyberanarchist efforts to bypass the state via blockchain technology, to the claims that Google and Facebook, as they mediate between billions of users, are relevant non-state actor in international affairs, to Lessig’s (1999) ever prescient claim that “Code is Law”, there is undoubtedly something going on with technology’s relationship to the state which is worth paying attention to.

There is an intellectual temptation (one that I myself am prone to) to take seriously the possibility of a fully autonomous technological alternative to the state. Something like a constitution written in source code has an appeal: it would be clear, precise, and presumably based on something like a consensus of those who participate in its creation. It is also an idea that can be frightening (Give up all control to the machines?) or ridiculous. The example of The DAO, the Ethereum ‘distributed autonomous organization’ that raised millions of dollars only to have them stolen in a technical hack, demonstrates the value of traditional legal institutions which protect the parties that enter contracts with processes that ensure fairness in their interpretation and enforcement.

It is more sociologically accurate, in any case, to consider software, hardware, and data collection not as autonomous actors but as parts of a sociotechnical system that maintains and modifies it. This is obvious to practitioners, who spend their lives negotiating the social systems that create technology. For those for whom it is not obvious, there’s reams of literature on the social embededness of “algorithms” (Gillespie, 2014; Kitchin, 2017). These themes are recited again in recent critical work on Artificial Intelligence; there are those that wisely point out that a functioning artificially intelligent system depends on a lot of labor (those who created and cleaned data, those who built the systems they are implemented on, those that monitor the system as it operates) (Kelkar, 2017). So rather than discussing the role of particular technologies as alternatives to the state, we should shift our focus to the great variety of sociotechnical organizations.

One thing that is apparent, when taking this view, is that states, as traditionally conceived, are themselves sociotechnical organizations. This is, again, an obvious point well illustrated in economic histories such as (Beniger, 1986). Communications infrastructure is necessary for the control and integration of society, let alone effective military logistics. The relationship between those industrial actors developing this infrastructure, whether it be building roads, running a postal service, laying rail or telegram wires, telephone wires, satellites, Internet protocols, and now social media–and the state has always been interesting and a story of great fortunes and shifts in power.

What is apparent after a serious look at this history is that political theory, especially liberal political theory as it developed in the 1700’s an onward as a theory of the relationship between individuals bound by social contract emerging from nature to develop a just state, leaves out essential scientific facts of the matter of how society has ever been governed. Control of communications and control infrastructure has never been equally dispersed and has always been a source of power. Late modern rearticulations of liberal theory and reactions against it (Rawls and Nozick, both) leave out technical constraints on the possibility of governance and even the constitution of the subject on which a theory of justice would have its ground.

Were political theory to begin from a more realistic foundation, it would need to acknowledge the existence of sociotechnical organizations as a political unit. There is a term for this view, “managerialism“, which, as far as I can tell is used somewhat pejoratively, like “neoliberalism”. As an “-ism”, it’s implied that managerialism is an ideology. When we talk about ideologies, what we are doing is looking from an external position onto an interdependent set of beliefs in their social context and identifying, through genealogical method or logical analysis, how those beliefs are symptoms of underlying causes that are not precisely as represented within those beliefs themselves. For example, one critiques neoliberal ideology, which purports that markets are the best way to allocate resources and advocates for the expansion of market logic into more domains of social and political life, but pointing out that markets are great for reallocating resources to capitalists, who bankroll neoliberal ideologues, but that many people who are subject to neoliberal policies do not benefit from them. While this is a bit of a parody of both neoliberalism and the critiques of it, you’ll catch my meaning.

We might avoid the pitfalls of an ideological managerialism (I’m not sure what those would be, exactly, having not read the critiques) by taking from it, to begin with, only the urgency of describing social reality in terms of organization and management without assuming any particular normative stake. It will be argued that this is not a neutral stance because to posit that there is organization, and that there is management, is to offend certain kinds of (mainly academic) thinkers. I get the sense that this offendedness is similar to the offense taken by certain critical scholars to the idea that there is such a thing as scientific knowledge, especially social scientific knowledge. Namely, it is an offense taken to the idea that a patently obvious fact entails ones own ignorance of otherwise very important expertise. This is encouraged by the institutional incentives of social science research. Social scientists are required to maintain an aura of expertise even when their particular sub-discipline excludes from its analysis the very systems of bureaucratic and technical management that its university depends on. University bureaucracies are, strangely, in the business of hiding their managerialist reality from their own faculty, as alternative avenues of research inquiry are of course compelling in their own right. When managerialism cannot be contested on epistemic grounds (because the bluff has been called), it can be rejected on aesthetic grounds: managerialism is not “interesting” to a discipline, perhaps because it does not engage with the personal and political motivations that constitute it.

What sets managerialism aside from other ideologies, however, is that when we examine its roots in social context, we do not discover a contradiction. Managerialism is not, as far as I can tell, successful as a popular ideology. Managerialism is attractive only to that rare segment of the population that work closely with bureaucratic management. It is here that the technical constraints of information flow and its potential uses, the limits of autonomy especially as it confronts the autonomies of others, the persistence of hierarchy despite the purported flattening of social relations, and so on become unavoidable features of life. And though one discovers in these situations plenty of managerial incompetence, one also comes to terms with why that incompetence is a necessary feature of the organizations that maintain it.

Little of what I am saying here is new, of course. It is only new in relation to more popular or appealing forms of criticism of the relationship between technology, organizations, power, and ethics. So often the political theory implicit in these critiques is a form of naive egalitarianism that sees a differential in power as an ethical red flag. Since technology can give organizations a lot of power, this generates a lot of heat around technology ethics. Starting from the perspective of an ethicist, one sees an uphill battle against an increasingly inscrutable and unaccountable sociotechnical apparatus. What I am proposing is that we look at things a different way. If we start from general principles about technology its role in organizations–the kinds of principles one would get from an analysis of microeconomic theory, artificial intelligence as a mathematical discipline, and so on–one can try to formulate managerial constraints that truly confront society. These constraints are part of how subjects are constituted and should inform what we see as “ethical”. If we can broker between these hard constraints and the societal values at stake, we might come up with a principle of justice that, if unpopular, may at least be realistic. This would be a contribution, at the end of the day, to political theory, not as an ideology, but as a philosophical advance.

References

Beniger, James R. “The Control Revolution: Technological and Economic Origins of the.” Information Society (1986).

Bird, Sarah, et al. “Exploring or Exploiting? Social and Ethical Implications of Autonomous Experimentation in AI.” (2016).

Gillespie, Tarleton. “The relevance of algorithms.” Media technologies: Essays on communication, materiality, and society 167 (2014).

Kelkar, Shreeharsh. “How (Not) to Talk about AI.” Platypus, 12 Apr. 2017, blog.castac.org/2017/04/how-not-to-talk-about-ai/.

Kitchin, Rob. “Thinking critically about and researching algorithms.” Information, Communication & Society 20.1 (2017): 14-29.

Lessig, Lawrence. “Code is law.” The Industry Standard 18 (1999).