







Open research data services have matured to the point where the cost of sustaining them at scale has become a primary design constraint, driving providers to make deliberate choices that may reduce user convenience to keep the service viable. The FAIR (Findable, Accessible, Interoperable, Reuseable) principles describe whether a dataset is well stewarded, and FAIR compliance is often treated as a proxy for usability. FAIR does not capture the cost to a user of finding, accessing, interpreting, and applying a dataset. We introduce the Dataset Friction Framework (DFF) as a complement to FAIR, directly addressing usability. DFF measures user-facing friction across six dimensions, distinguishing engineered friction (deliberate data provider design choices that sustain a service) from accidental friction (defects that require remediation). The framework is validated against 18,556 support tickets from the European Centre for Medium-Range Weather Forecasts (January 2024 to May 2026), which serves 280,000 registered users. Restricting the analysis to tickets raised by external reporters reduces the corpus by 12.3%, but every dimension's internal-staff share falls below this baseline -- confirming that the reported friction signals are genuinely user-facing. We then assess three real datasets across three providers and show that FAIR compliance and DFF friction can disagree in both directions: a 92% FAIR-compliant dataset can still carry substantial friction, and a 42% FAIR score can be an artefact of anti-scraping policy rather than poor stewardship. The two measures are non-redundant and jointly informative: FAIR compliance does not predict DFF friction in either direction. This constitutes the first large-scale empirical application of the framework; cross-institutional validation is identified as the immediate next step.
OpenDataVal: a Unified Benchmark for Data Valuation
Assessing the quality and impact of individual data points is critical for improving model performance and mitigating undesirable biases within the training dataset. Several data valuation algorithms have been proposed to quantify data quality, however, there lacks a systemic and standardized benchmarking system for data valuation. In this paper, we introduce OpenDataVal, an easy-to-use and unified benchmark framework that empowers researchers and practitioners to apply and compare various data valuation algorithms. OpenDataVal provides an integrated environment that includes (i) a diverse collection of image, natural language, and tabular datasets, (ii) implementations of eleven different state-of-the-art data valuation algorithms, and (iii) a prediction model API that can import any models in scikit-learn. Furthermore, we propose four downstream machine learning tasks for evaluating the quality of data values. We perform benchmarking analysis using OpenDataVal, quantifying and comparing the efficacy of state-of-the-art data valuation approaches. We find that no single algorithm performs uniformly best across all tasks, and an appropriate algorithm should be employed for a user's downstream task. OpenDataVal is publicly available at https://opendataval.github.io with comprehensive documentation. Furthermore, we provide a leaderboard where researchers can evaluate the effectiveness of their own data valuation algorithms.
FAIR Principles - GO FAIR
In 2016, the ‘FAIR Guiding Principles for scientific data management and stewardship’ were published in Scientific Data. The authors intended to provide guidelines to improve the Findability, Accessibility, Interoperability, and Reuse of digital assets. The principles emphasise machine-actionability (i.e., the capacity of… Continue reading →

A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

2D-Shapley: A Framework for Fragmented Data Valuation
Data valuation—quantifying the contribution of individual data sources to certain predictive behaviors of a model—is of great importance to enhancing the transparency of machine learning and designing incentive systems for data sharing. Existing work has focused on evaluating data sources with the shared feature or sample space. How to valuate fragmented data sources of which each only contains partial features and samples remains an open question. We start by presenting a method to calculate the counterfactual of removing a fragment from the aggregated data matrix. Based on the counterfactual calculation, we further propose 2D-Shapley, a theoretical framework for fragmented data valuation that uniquely satisfies some appealing axioms in the fragmented data context. 2D-Shapley empowers a range of new use cases, such as selecting useful data fragments, providing interpretation for sample-wise data values, and fine-grained data issue diagnosis.
The case against efficiency: friction in social media
Social media platforms frequently prioritize efficiency to maximize ad revenue and user engagement, often sacrificing deliberation, trust, and reflective, purposeful cognitive engagement in the process. This manuscript examines the potential of friction—design choices that intentionally slow user interactions—as an alternate approach. We present a case against efficiency as the dominant paradigm on social media and advocate for a complex systems approach to understanding and analyzing friction. Drawing from interdisciplinary literature, real-world examples, and industry experiments, we highlight the potential for friction to mitigate issues like polarization, disinformation, and toxic content without resorting to censorship. We propose a state space representation of friction to establish a multidimensional framework and language for analyzing the diverse forms and functions through which friction can be implemented. Additionally, we propose several experimental designs to examine the impact of friction on system dynamics, user behavior, and information ecosystems, each designed with complex systems solutions and perspectives in mind. Our case against efficiency underscores the critical role of friction in shaping digital spaces, challenging the relentless pursuit of efficiency and exploring the potential of thoughtful slowing.

Benchmarking World-Model Learning with Environment-Level Queries
World models are central to building AI agents capable of flexible reasoning and planning. Yet current evaluations (i) test only properties measurable from observed interactions, such as next-frame prediction or task return, and (ii) do not test whether a learned model supports diverse queries about the environment. In contrast, humans build $\textit{general-purpose}$ models that can answer many different questions about an environment$\unicode{x2014}$including questions that require understanding global structure and counterfactual consequences. We propose $\textit{WorldTest}$: a protocol for evaluating whether agents learn models that support multiple $\textit{environment-level queries}\unicode{x2014}$questions whose answers depend on properties of the full environment, not just observed trajectories. Individually, these queries can target properties (e.g., reachability or the effects of interventions) that no single rollout distribution determines. Collectively, they assess model generality across query types. We instantiate WorldTest as $\textit{AutumnBench}$, a benchmark of 43 interactive grid-world environments and 129 tasks across three query families for both humans and learning agents. Experiments with 517 human participants and five frontier models show that humans substantially outperform these models, a gap we attribute to differences in exploration and belief updating. AutumnBench provides a framework for evaluating world-model learning in grid-world environments with environment-level queries, and WorldTest provides a template for extending such evaluations to richer domains.

Adaptive Drift Defense: A Unified Framework for Data, Task, And User-Intent Drift in LLM Apps
The deployment of Large Language Models (LLMs) to production systems where user expectations, data source and task requirements are constantly changing is increasing. This change brings in the element of drift--changes to input distributions, tool-call patterns or user intent--that impairs performance given enough time and is passed unnoticed. Current approaches to drift management tend to be either excessively specific to data-level drift monitoring or directly retrain the model which represents an unacceptable resource-intensive task; thus, much ground remains to be lost with regards to real-time response to drift and resource requirement. In this paper we introduce coherent framework Adaptive Drift Defense, which can integrate three orthogonal layers of detection, like retrieval distribution monitoring, tool-call graph analysis or estimation of output variance, to detect data, task and user-intent drift in a concurrent way. A computing bandit mitigation policy actively chooses timely refinements, retrieval adaptations or tool-routing policies, driving performance to equilibrium in terms of requiring retraining of the model. Experiments on major customer care helpers (~1.2M interactions) and business analytics copilots (~800k queries) show that the system achieves an 88 percent accuracy and 86 percent recall in detecting drift, with 20-35 percent savings in the cost of manual rework with insignificant latency overhead (<50ms). When compared to non-incremental pipelines, task success rates were increased by 15-20 percent, bridging most of the performance-gap to full retraining with only 12-percent extra compute cost. These results imply the conceptual feasibility of conventional, collaborative drift monitoring of LLM system. Analytical paper also furnishes reference dashboards and operational playbooks which assists deployment teams. These results help us understand that adaptive methods with mitigation-first approaches will enable high-quality service quality in highly dynamic settings and that it scales well to situations of frequent model retraining.
Data Distribution Valuation
Data valuation is a class of techniques for quantitatively assessing the value of data for applications like pricing in data marketplaces. Existing data valuation methods define a value for a discrete dataset. However, in many use cases, users are interested in not only the value of the dataset, but that of the distribution from which the dataset was sampled. For example, consider a buyer trying to evaluate whether to purchase data from different vendors. The buyer may observe (and compare) only a small preview sample from each vendor, to decide which vendor's data distribution is most useful to the buyer and purchase. The core question is how should we compare the values of data distributions from their samples? Under a Huber characterization of the data heterogeneity across vendors, we propose a maximum mean discrepancy (MMD)-based valuation method which enables theoretically principled and actionable policies for comparing data distributions from samples. We empirically demonstrate that our method is sample-efficient and effective in identifying valuable data distributions against several existing baselines, on multiple real-world datasets (e.g., network intrusion detection, credit card fraud detection) and downstream applications (classification, regression).
Promoting User Data Autonomy During the Dissolution of a Monopolistic Firm
The deployment of AI in consumer products is currently focused on the use of so-called foundation models, large neural networks pre-trained on massive corpora of digital records. This emphasis on scaling up datasets and pre-training computation raises the risk of further consolidating the industry, and enabling monopolistic (or oligopolistic) behavior. Judges and regulators seeking to improve market competition may employ various remedies. This paper explores dissolution -- the breaking up of a monopolistic entity into smaller firms -- as one such remedy, focusing in particular on the technical challenges and opportunities involved in the breaking up of large models and datasets. We show how the framework of Conscious Data Contribution can enable user autonomy during under dissolution. Through a simulation study, we explore how fine-tuning and the phenomenon of "catastrophic forgetting" could actually prove beneficial as a type of machine unlearning that allows users to specify which data they want used for what purposes.

Learning to Limit Data Collection via Scaling Laws: A Computational Interpretation for the Legal Principle of Data Minimization
Modern machine learning systems are increasingly characterized by extensive personal data collection, despite the diminishing returns and increasing societal costs of such practices. Yet, data minimisation is one of the core data protection principles enshrined in the European Union's General Data Protection Regulation ('GDPR') and requires that only personal data that is adequate, relevant and limited to what is necessary is processed. However, the principle has seen limited adoption due to the lack of technical interpretation. In this work, we build on literature in machine learning and law to propose FIDO, a Framework for Inhibiting Data Overcollection. FIDO learns to limit data collection based on an interpretation of data minimization tied to system performance. Concretely, FIDO provides a data collection stopping criterion by iteratively updating an estimate of the performance curve, or the relationship between dataset size and performance, as data is acquired. FIDO estimates the performance curve via a piecewise power law technique that models distinct phases of an algorithm's performance throughout data collection separately. Empirical experiments show that the framework produces accurate performance curves and data collection stopping criteria across datasets and feature acquisition algorithms. We further demonstrate that many other families of curves systematically overestimate the return on additional data. Results and analysis from our investigation offer deeper insights into the relevant considerations when designing a data minimization framework, including the impacts of active feature acquisition on individual users and the feasability of user-specific data minimization. We conclude with practical recommendations for the implementation of data minimization.

Democratizing Data
Democratizing Data builds a community-driven data ecosystem by identifying how datasets are used and reducing barriers to accessing high-quality public data. The initiative enhances the discoverability, usability, and relevance of data for researchers, policymakers, and stakeholder communities. A suite of tools and strategic partnerships supports this work by connecting users to the data, insights, and networks needed to inform decisions and generate impact.
Distributionally Robust Data Valuation
Data valuation quantifies the contribution of each data point to the performance of a machine learning model. Existing works typically define the value of data by its improvement of the validation performance of the trained model. However, this approach can be impractical to apply in collaborative machine learning and data marketplace since it is difficult for the parties/buyers to agree on a common validation dataset or determine the exact validation distribution a priori. To address this, we propose a distributionally robust data valuation approach to perform data valuation without known/fixed validation distributions. Our approach defines the value of data by its improvement of the distributionally robust generalization error (DRGE), thus providing a worst-case performance guarantee without a known/fixed validation distribution. However, since computing DRGE directly is infeasible, we propose using model deviation as a proxy for the marginal improvement of DRGE (for kernel regression and neural networks) to compute data values. Furthermore, we identify a notion of uniqueness where low uniqueness characterizes low-value data. We empirically demonstrate that our approach outperforms existing data valuation approaches in data selection and data removal tasks on real-world datasets (e.g., housing price prediction, diabetes hospitalization prediction).
Perspective Chapter: Fit for Purpose? Creative Commons Licensing for Research Data in the Age of Artificial Intelligence
Licensing is an important component of the re-usability of research data, itself part of the FAIR principles: without clear, machine-readable licensing, datasets risk becoming technically...

The Friction is Your Judgment — Armin Ronacher & Cristina Poncela Cubeiro, Earendil
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking progress, and identifying weaknesses in foundation and non-foundation models. They can inform model selection for downstream tasks and influence policy initiatives. However, not all benchmarks are the same: their quality depends on their design and usability. In this paper, we develop an assessment framework considering 46 best practices across an AI benchmark's lifecycle and evaluate 24 AI benchmarks against it. We find that there exist large quality differences and that commonly used benchmarks suffer from significant issues. We further find that most benchmarks do not report statistical significance of their results nor allow for their results to be easily replicated. To support benchmark developers in aligning with best practices, we provide a checklist for minimum quality assurance based on our assessment. We also develop a living repository of benchmark assessments to support benchmark comparability, accessible at betterbench.stanford.edu.

The Challenge of Understanding What Users Want: Inconsistent Preferences and Engagement Optimization
Online platforms have a wealth of data, run countless experiments, and use industrial-scale algorithms to optimize user experience. Despite this, many users seem to regret the time they spend on these platforms. One possible explanation is that incentives are misaligned: platforms are not optimizing for user happiness. We suggest the problem runs deeper, transcending the specific incentives of any particular platform, and instead stems from a mistaken foundational assumption. To understand what users want, platforms look at what users do. This is a kind of revealed-preference assumption that is ubiquitous in the way user models are built. Yet research has demonstrated, and personal experience affirms, that we often make choices in the moment that are inconsistent with what we actually want. The behavioral economics and psychology literatures suggest, for example, that we can choose mindlessly or that we can be too myopic in our choices, behaviors that feel entirely familiar on online platforms. In this work, we develop a model of media consumption where users have inconsistent preferences. We consider a platform which wants to maximize user utility, but only observes behavioral data in the form of the user’s engagement. We show how our model of users’ preference inconsistencies produces phenomena that are familiar from everyday experience but difficult to capture in traditional user interaction models. These phenomena include users who have long sessions on a platform but derive very little utility from it, and platform changes that steadily raise user engagement before abruptly causing users to go “cold turkey” and quit. A key ingredient in our model is a formulation for how platforms determine what to show users: they optimize over a large set of potential content (the content manifold) parametrized by underlying features of the content. Whether improving engagement improves user welfare depends on the direction of movement in the content manifold: For certain directions of change, increasing engagement makes users less happy, whereas in other directions on the same manifold, increasing engagement makes users happier. We provide a characterization of the structure of content manifolds for which increasing engagement fails to increase user utility. By linking these effects to abstractions of platform design choices, our model thus creates a theoretical framework and vocabulary in which to explore interactions between design, behavioral science, and social media. This paper was accepted by Yan Chen, behavioral economics and decision analysis. Funding: This work was supported by the Vannevar Bush Faculty Fellowship and Multidisciplinary University Research Initiative [Grant W911NF-19-0217]. Supplemental Material: The online appendices are available at https://doi.org/10.1287/mnsc.2022.03683 .
