







Data science is concerned with finding answers to questions on the basis of available data, and communicating that effort. Besides showing the results, this communication involves sharing the data used, but also exposing the path that led to the answers in a comprehensive and reproducible way. It also acknowledges the fact that available data may not be sufficient to answer questions, and that any answers are conditional on the data collection or sampling protocols employed.
A Data Utopia for Science-of-Science
Here I want to briefly sketch out a vision for how to solve a key set of problems facing science-of-science researchers, using the relatively new idea of a ‘data trust.’ In my ideal wor…

Cambria | Proceedings of the 8th Workshop on Principles and Practice of Consistency for Distributed Data
This summary was generated using automated tools and was not authored or reviewed by the article's author(s). It is provided to support discovery, help readers assess relevance, and assist readers from adjacent research areas in understanding the work. It is intended to complement the author-supplied abstract, which remains the primary summary of the paper. The full article remains the authoritative version of record. Click here to learn more.
Dynamic Data Science
<p>Use our Common Online Data Analysis Platform (CODAP) to explore dynamic data science activities and gain fluency in data moves to examine large datasets.</p>


Related work | Data Counterfactuals
Adjacent research areas and starting papers, curated in Semble and synced into the site.

What makes something data?
This is a question I posted on BlueSky on Friday 11/21/25, inspired by a talk I recently attended about evaluation of “AI” systems. I think…

datasetpapers — a public research experiment
An experimental approach to versioned, forkable, machine-readable analyses. A prototype, not a product or service.

datasetpapers — a public research experiment
An experimental approach to versioned, forkable, machine-readable analyses. A prototype, not a product or service.

Proposal: User Intents for Data Reuse · bluesky-social atproto · Discussion #3617
This is a discussion thread for the User Intents for Data Reuse proposal.
Data Integration
This book is an introduction to the problem of data integration and a rigorous account of one of the leading approaches to solving this problem.

Data Feminism
Today, data science is a form of power. It has been used to expose injustice, improve health outcomes, and topple governments. But it has also been used to d...

Democratizing Data
Democratizing Data builds a community-driven data ecosystem by identifying how datasets are used and reducing barriers to accessing high-quality public data. The initiative enhances the discoverability, usability, and relevance of data for researchers, policymakers, and stakeholder communities. A suite of tools and strategic partnerships supports this work by connecting users to the data, insights, and networks needed to inform decisions and generate impact.
Datacurve | The data engine for frontier AI
Custom data for long-horizon reasoning, software engineering, and data science.

Part of the datacounterfactuals.org reading lists. Research on data provenance, dataset documentation, licensing and attribution audits, and technical source-attribution methods for understanding which data sources are available, permitted, or responsible for model behavior.
WASA: WAtermark-based Source Attribution for Large Language Model-Generated Data
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

Datasheets for Datasets
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

A large-scale audit of dataset licensing and attribution in AI