







Human society is coordinated by mechanisms that control how prices are agreed, taxes are set, and electoral votes are tallied. The design of robust...
Machine learning political orders
A significant set of epistemic and political transformations are taking place as states and societies begin to understand themselves and their problems through the paradigm of deep neural network algorithms. A machine learning political order does not merely change the political technologies of governance, but is itself a reordering of politics, of what the political can be. When algorithmic systems reduce the pluridimensionality of politics to the output of a model, they simultaneously foreclose the potential for other political claims to be made and alternative political projects to be built. More than this foreclosure, a machine learning political order actively profits and learns from the fracturing of communities and the destabilising of democratic rights. The transformation from rules-based algorithms to deep learning models has paralleled the undoing of rules-based social and international orders – from the use of machine learning in the campaigns of the UK EU referendum, to the trialling of algorithmic immigration and welfare systems, and the use of deep learning in the COVID-19 pandemic – with political problems becoming reconfigured as machine learning problems. Machine learning political orders decouple their attributes, features and clusters from underlying social values, no longer tethered to notions of good governance or a good society, but searching instead for the optimal function of abstract representations of data.

Sharing the Algorithm: The Tax Solution to Generative AI
This article argues that tax policy offers a core tool for mitigating the sweeping public policy challenges of generative Artificial Intelligence ("AI"
AI Behavioral Science
We outline a foundation for a new field of ``AI Behavioral Science,'' covering three perspectives. First, as AI becomes ubiquitous and is increasingly proprietary and opaque, it becomes vital to develop techniques for assessing AI behavior. We outline how tools developed to assess people's behaviors by social scientists can be used to assess and infer AI's behaviors biases, tendencies, and heuristics. Second, we also discuss how AI can change the ways in which we learn about human behavior. Beyond its computational power, AI offers new techniques for simulating, inferring, and predicting human behaviors that we outline and discuss. Third, as humans and AI are interacting in increasingly complex and intertwined systems, we need to understand the implications for the resulting economic and political outcomes. We outline issues that are increasingly pressing concerning the future of human-AI interactions and potential changes and disruptions that can ensue.

The GenAI Future of Consumer Research
Abstract We develop a novel generative AI (GenAI) trajectory, “democratization-average trap-model collapse,” to identify data and model challenges posed by GenAI, from which we project the GenAI future of consumer research. This trajectory consists of three key phenomena: democratization broadens consumer participation, the average trap produces generic responses, and model collapse occurs when GenAI outputs lose human sensibilities. Data and model challenges arise as democratization enhances data representation while also embedding real-world biases. The average trap, caused by next-token prediction models, leads to generic outputs that lack individuality. Additionally, model collapse occurs when GenAI increasingly learns from its own outputs, amplifying machine bias and diverging from human behavior. To address these challenges, researchers can leverage democratization to study marginalized consumers and prioritize human-centered research over purely data-driven methods. The average trap can be mitigated by fine-tuning models with task-specific and marginalized consumption data while engineering responses for uniqueness. Preventing model collapse requires integrating human–machine hybrid data and applying theories of mind to realign AI with human-centric consumption. Finally, we outline three future research directions: preserving data distribution tails to support consumption democratization, countering the average trap in next-token prediction, and reversing the trajectory from democratization to model collapse.

A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments
Environments built for people are increasingly operated by a new class of economic actors: LLM-powered software agents making decisions on our behalf. These decisions range from our purchases to travel plans to medical treatment selection. Current evaluations of these agents largely focus on task competence, but we argue for a deeper assessment: how these agents choose when faced with realistic decisions. We introduce ABxLab, a framework for systematically probing agentic choice through controlled manipulations of option attributes and persuasive cues. We apply this to a realistic web-based shopping environment, where we vary prices, ratings, and psychological nudges, all of which are factors long known to shape human choice. We find that agent decisions shift predictably and substantially in response, revealing that agents are strongly biased choosers even without being subject to the cognitive constraints that shape human biases. This susceptibility reveals both risk and opportunity: risk, because agentic consumers may inherit and amplify human biases; opportunity, because consumer choice provides a powerful testbed for a behavioral science of AI agents, just as it has for the study of human behavior. We release our framework as an open benchmark for rigorous, scalable evaluation of agent decision-making.

Society-in-the-loop: programming the algorithmic social contract
Recent rapid advances in Artificial Intelligence (AI) and Machine Learning have raised many questions about the regulatory and governance mechanisms for autonomous machines. Many commentators, scholars, and policy-makers now call for ensuring that algorithms governing our lives are transparent, fair, and accountable. Here, I propose a conceptual framework for the regulation of AI and algorithmic systems. I argue that we need tools to program, debug and maintain an algorithmic social contract, a pact between various human stakeholders, mediated by machines. To achieve this, we can adapt the concept of human-in-the-loop (HITL) from the fields of modeling and simulation, and interactive machine learning. In particular, I propose an agenda I call society-in-the-loop (SITL), which combines the HITL control paradigm with mechanisms for negotiating the values of various stakeholders affected by AI systems, and monitoring compliance with the agreement. In short, ‘SITL = HITL + Social Contract.’

Reinforcement Learning from Human Feedback
The authoritative guide for Reinforcement learning from human feedback, alignment, and post-training LLMs. Aligning AI models to human preferences helps them become safer, smarter, easier to use, and tuned to the exact style the creator desires. Reinforcement Learning From Human Feedback (RHLF) is the process for using human responses to a model’s output to shape its alignment, and therefore its behavior. In Reinforcement Learning from Human Feedback, author Nathan Lambert blends diverse perspectives from fields like philosophy and economics with the core mathematics and computer science of RLHF to provide a practical guide you can use to apply RLHF to your models. In Reinforcement Learning from Human Feedback you’ll discover: How today’s most advanced AI models are taught from human feedback How large-scale preference data is collected and how to improve your data pipelines A comprehensive overview with derivations and implementations for the core policy-gradient methods used to train AI models with reinforcement learning (RL) Direct Preference Optimization (DPO), direct alignment algorithms, and simpler methods for preference finetuning How RLHF methods led to the current reinforcement learning from verifiable rewards (RLVR) renaissance Tricks used in industry to round out models, from product, character or personality training, AI feedback, and more How to approach evaluation and how evaluation has changed over the years Standard recipes for post-training combining more methods like instruction tuning with RLHF Behind-the-scenes stories from building open models like Llama-Instruct, Zephyr, Olmo, and Tülu After ChatGPT used RLHF to become production-ready, this foundational technique exploded in popularity. In Reinforcement Learning from Human Feedback, AI expert Nathan Lambert gives a true industry insider's perspective on modern RLHF training pipelines, and their trade-offs. Using hands-on experiments and mini-implementations, Nathan clearly and concisely introduces the alignment techniques that can transform a generic base model into a human-friendly tool.

Artificial Intelligence, Algorithmic Pricing, and Collusion
Increasingly, algorithms are supplanting human decision-makers in pricing goods and services. To analyze the possible consequences, we study experimentally the behavior of algorithms powered by Artificial Intelligence (Q-learning) in a workhorse oligopoly model of repeated price competition. We find that the algorithms consistently learn to charge supracompetitive prices, without communicating with one another. The high prices are sustained by collusive strategies with a finite phase of punishment followed by a gradual return to cooperation. This finding is robust to asymmetries in cost or demand, changes in the number of players, and various forms of uncertainty.
Simple Pricing | Machine Learning Infrastructure | Deep Infra
We provide only pay-what-you-use pricing with no long-term contracts or upfront costs for our machine learning models and infrastructure. Learn more!

From Predictive Algorithms to Automatic Generation of Anomalies
Machine learning algorithms can find predictive signals that researchers fail to notice; yet they are notoriously hard-to-interpret. How can we extract theoretical insights from these black boxes? History provides a clue. Facing a similar problem – how to extract theoretical insights from their intuitions – researchers often turned to “anomalies:” constructed examples that highlight flaws in an existing theory and spur the development of new ones. Canonical examples include the Allais paradox and the Kahneman-Tversky choice experiments for expected utility theory. We suggest anomalies can extract theoretical insights from black box predictive algorithms. We develop procedures to automatically generate anomalies for an existing theory when given a predictive algorithm. We cast anomaly generation as an adversarial game between a theory and a falsifier, the solutions to which are anomalies: instances where the black box algorithm predicts - were we to collect data - we would likely observe violations of the theory. As an illustration, we generate anomalies for expected utility theory using a large, publicly available dataset on real lottery choices. Based on an estimated neural network that predicts lottery choices, our procedures recover known anomalies and discover new ones for expected utility theory. In incentivized experiments, subjects violate expected utility theory on these algorithmically generated anomalies; moreover, the violation rates are similar to observed rates for the Allais paradox and Common ratio effect.

Data Shapley: Equitable Valuation of Data for Machine Learning
As data becomes the fuel driving technological and economic growth, a fundamental challenge is how to quantify the value of data in algorithmic predictions and decisions. For example, in healthcare and consumer markets, it has been suggested that individuals should be compensated for the data that they generate, but it is not clear what is an equitable valuation for individual data. In this work, we develop a principled framework to address data valuation in the context of supervised machine learning. Given a learning algorithm trained on $n$ data points to produce a predictor, we propose data Shapley as a metric to quantify the value of each training datum to the predictor performance. Data Shapley uniquely satisfies several natural properties of equitable data valuation. We develop Monte Carlo and gradient-based methods to efficiently estimate data Shapley values in practical settings where complex learning algorithms, including neural networks, are trained on large datasets. In addition to being equitable, extensive experiments across biomedical, image and synthetic data demonstrate that data Shapley has several other benefits: 1) it is more powerful than the popular leave-one-out or leverage score in providing insight on what data is more valuable for a given learning task; 2) low Shapley value data effectively capture outliers and corruptions; 3) high Shapley value data inform what type of new data to acquire to improve the predictor.
Integrative experiments identify how punishment affects welfare in public goods games
Despite decades of research, the conditions under which punishment promotes cooperation remain unclear. Through an integrative experiment varying 14 design parameters of public goods games across 360 experimental conditions (147,618 decisions from 7100 participants), we reveal substantial heterogeneity in punishment effectiveness: Its impact on welfare ranges from 43% improvement to 44% reduction depending on the game parameters. To characterize these patterns, we developed models that outperformed human forecasters in predicting punishment effectiveness in new experiments. Communication emerges as the most important factor, followed by contribution framing (opt out versus opt in), contribution type (variable versus all-or-nothing), game length, and outcome visibility, though these factors often interact. The results reframe the debate from whether punishment works to when it does, demonstrating how integrative experiments enable discovery of generalizable patterns in social phenomena. , Editor’s summary People face conflicts between maximizing personal gain versus supporting collective interests. If we cooperatively recycle or donate to charities, it benefits society, but it also costs us time and resources that could be selfishly preserved for ourselves. We impose penalties to deter those undesirable or selfish behaviors, but under what conditions do punishments or penalties effectively modify behavior to benefit group welfare? Alsobay et al . systematically and simultaneously varied 14 factors together instead of in isolation. Punishment was unequivocally most effective when paired with consistent communication, particularly over time. Another effective factor was “opting out” or withdrawing some, but not all, endowments already in the public fund. These methodological advances revealed when, rather than whether, punishment works. —Ekeoma Uzogara , INTRODUCTION Human societies face many situations where individual and collective interests conflict, often referred to as social dilemmas. Costly peer punishment has been studied for more than 25 years in public goods games (stylized behavioral experiments in which individuals decide how much to contribute to a shared pool that benefits everyone) as a mechanism to promote cooperation. Prior research has identified many contextual factors that moderate punishment’s effectiveness, including game length, communication, group size, punishment cost, and so on. However, the specific conditions under which punishment improves group welfare remain unclear. RATIONALE We argue that this lack of clarity derives from the dominant experimental paradigm, in which any given study manipulates only one or a few theoretically informed factors. Because such studies differ in many ways (different experimental procedures, populations), their results are often difficult to compare or integrate. Consequently, one can list many factors that have some effect, but cannot say how much each matters relative to the others, or how they work together, and as a result, cannot predict when punishment will help or harm welfare in new settings. To address this fundamental knowledge gap, we use an integrative experimental design and systematically vary 14 parameters across 360 conditions (147,618 decisions from 7100 participants) to elucidate when punishment improves versus undermines welfare in public goods games, which factors matter most, and how they interact. RESULTS The effect of punishment on welfare ranged from 43% improvement to 44% reduction depending on the specific combination of game parameters. To characterize this heterogeneity, we trained a model that outperformed all 553 human forecasters (laypeople and experts) in predicting whether punishment would help or harm welfare in new experiments. Communication emerged as roughly three times more important than any other factor, followed by contribution framing (opt in versus opt out), contribution type (variable versus all-or-nothing), game length, and peer outcome visibility (whether participants can see others’ earnings). These factors often interact. For example, longer games enhance punishment’s effectiveness only when communication is available, and contribution framing effects depend on both contribution type and outcome visibility. CONCLUSION Many phenomena in social science are shaped by many factors whose interactions are consequential, yet the dominant experimental paradigm often limits its inquiry to “does a given effect exist?” and examines hypothesized factors in isolation. As a result, research programs can accumulate many partial explanations without a clear picture of how they combine to determine outcomes across settings. Knowing that factors matter individually is fundamentally different from knowing how much each matters and how they interact. The integrative approach implemented here offers one way forward. It varies many factors simultaneously within a shared design space, evaluates models by their predictive accuracy on new experiments, and probes those models to constrain and develop theory. Our hope is that integrative experiment designs, combined with models that integrate prediction and explanation, represent a path toward more cumulative social science. Integrative experiment reveals when punishment helps versus harms. We systematically varied 14 design parameters across 360 experimental conditions. The effect of punishment on cooperation efficiency ranged from −44% to +43% depending on the specific game parameters. Communication emerged as three times more important than any other factor, followed by contribution framing, contribution type, and game length.

Algorithmic Collective Action in Machine Learning
We initiate a principled study of algorithmic collective action on digital platforms that deploy machine learning algorithms. We propose a simple theoretical model of a collective interacting with a firm's learning algorithm. The collective pools the data of participating individuals and executes an algorithmic strategy by instructing participants how to modify their own data to achieve a collective goal. We investigate the consequences of this model in three fundamental learning-theoretic settings: the case of a nonparametric optimal learning algorithm, a parametric risk minimizer, and gradient-based optimization. In each setting, we come up with coordinated algorithmic strategies and characterize natural success criteria as a function of the collective's size. Complementing our theory, we conduct systematic experiments on a skill classification task involving tens of thousands of resumes from a gig platform for freelancers. Through more than two thousand model training runs of a BERT-like language model, we see a striking correspondence emerge between our empirical observations and the predictions made by our theory. Taken together, our theory and experiments broadly support the conclusion that algorithmic collectives of exceedingly small fractional size can exert significant control over a platform's learning algorithm.

Political Neutrality in AI Is Impossible- But Here Is How to Approximate It
AI systems often exhibit political bias, influencing users' opinions and decisions. While political neutrality-defined as the absence of bias-is often seen as an ideal solution for fairness and safety, this position paper argues that true political neutrality is neither feasible nor universally desirable due to its subjective nature and the biases inherent in AI training data, algorithms, and user interactions. However, inspired by Joseph Raz's philosophical insight that "neutrality [...] can be a matter of degree" (Raz, 1986), we argue that striving for some neutrality remains essential for promoting balanced AI interactions and mitigating user manipulation. Therefore, we use the term "approximation" of political neutrality to shift the focus from unattainable absolutes to achievable, practical proxies. We propose eight techniques for approximating neutrality across three levels of conceptualizing AI, examining their trade-offs and implementation strategies. In addition, we explore two concrete applications of these approximations to illustrate their practicality. Finally, we assess our framework on current large language models (LLMs) at the output level, providing a demonstration of how it can be evaluated. This work seeks to advance nuanced discussions of political neutrality in AI and promote the development of responsible, aligned language models.

Political Neutrality in AI Is Impossible- But Here Is How to Approximate It
AI systems often exhibit political bias, influencing users' opinions and decisions. While political neutrality-defined as the absence of bias-is often seen as an ideal solution for fairness and safety, this position paper argues that true political neutrality is neither feasible nor universally desirable due to its subjective nature and the biases inherent in AI training data, algorithms, and user interactions. However, inspired by Joseph Raz's philosophical insight that "neutrality [...] can be a matter of degree" (Raz, 1986), we argue that striving for some neutrality remains essential for promoting balanced AI interactions and mitigating user manipulation. Therefore, we use the term "approximation" of political neutrality to shift the focus from unattainable absolutes to achievable, practical proxies. We propose eight techniques for approximating neutrality across three levels of conceptualizing AI, examining their trade-offs and implementation strategies. In addition, we explore two concrete applications of these approximations to illustrate their practicality. Finally, we assess our framework on current large language models (LLMs) at the output level, providing a demonstration of how it can be evaluated. This work seeks to advance nuanced discussions of political neutrality in AI and promote the development of responsible, aligned language models.

Promoting User Data Autonomy During the Dissolution of a Monopolistic Firm
The deployment of AI in consumer products is currently focused on the use of so-called foundation models, large neural networks pre-trained on massive corpora of digital records. This emphasis on scaling up datasets and pre-training computation raises the risk of further consolidating the industry, and enabling monopolistic (or oligopolistic) behavior. Judges and regulators seeking to improve market competition may employ various remedies. This paper explores dissolution -- the breaking up of a monopolistic entity into smaller firms -- as one such remedy, focusing in particular on the technical challenges and opportunities involved in the breaking up of large models and datasets. We show how the framework of Conscious Data Contribution can enable user autonomy during under dissolution. Through a simulation study, we explore how fine-tuning and the phenomenon of "catastrophic forgetting" could actually prove beneficial as a type of machine unlearning that allows users to specify which data they want used for what purposes.
