







tfw @aaronstevenwhite.io brings an analysis as sharp as a knife to your half-baked Saturday-morning thoughts: aaronstevenwhite.leaflet.pub/3miwsz2hdv22i 🤯 If we're going to own our data, let's actually own our data. Which is to say: No, really, y'all, we're doing this. 💖🧠
Machine-readable attitudes - Computational Semantics++
aaronstevenwhite.leaflet.pubApr 7, 2026 at 10:57 PM
Permissioned Data Diary 4: The Big Picture - Daniel's Leaflets
A special edition of the data diary that sketches out the rough shape of where we're heading.
Living in Data: A Citizen's Guide to a Better Information Future (Paperback)
Jer Thorp’s analysis of the word “data” in 10,325 New York Times stories written between 1984 and 2018 shows a distinct trend: among the words most closely associated with “data,” we find not only its classic companions “information” and “digital,” but also a variety of new neighbors—from “scandal” and “misinformation” to “ethics,” “friends,” and “play.”To live in data in the twenty-first century

Who Even Cares About Data Ownership Anyway? - What The Function!?
Taking a look at a concept that's been talked a lot over the last few years, especially in the context of social media

The new way we’ll do science
Papers should become human-readable views over a graph of data, tools, results, and certificates.

The Consensus Trap: Dissecting Subjectivity and the “Ground Truth” Illusion in Data Annotation
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.


datasetpapers — a public research experiment
An experimental approach to versioned, forkable, machine-readable analyses. A prototype, not a product or service.

datasetpapers — a public research experiment
An experimental approach to versioned, forkable, machine-readable analyses. A prototype, not a product or service.

B-Sides: Permissioned data is a love triangle - Nick's Blog
This is a B-sides post with unpolished thoughts that didn’t make it into the main article.
Should We Treat Data as Labor? Moving Beyond “Free”
Should We Treat Data as Labor? Moving beyond "Free" by Imanol Arrieta-Ibarra, Leonard Goff, Diego Jiménez-Hernández, Jaron Lanier and E. Glen Weyl. Published in volume 108, pages 38-42 of AEA Papers and Proceedings, May 2018, Abstract: In the digital economy, user data is typically treated as capi...
Data Feminism
A new way of thinking about data science and data ethics that is informed by the ideas of intersectional feminism.

The Limits of Data
Policymakers want to make decisions based on clear data, but important factors are lost when we rely solely on data. A philosopher writes:

Can “Conscious Data Contribution” Help Users to Exert “Data Leverage” Against Technology Companies?
Tech users currently have limited ability to act on concerns regarding the negative societal impacts of large tech companies. However, recent work suggests that users can exert leverage using their role in the generation of valuable data, for instance by withholding their data contributions to intelligent technologies. We propose and evaluate a new means to exert this type of leverage against tech companies: "conscious data contribution" (CDC).
Nonrivalry and the Economics of Data
(September 2020) - Data is nonrival: a person's location history, medical records, and driving data can be used by many firms simultaneously. Nonrivalry leads to increasing returns. As a result, there may be social gains to data being used broadly across firms, even in the presence of privacy considerations. Fearing creative destruction, firms may choose to hoard their data, leading to the inefficient use of nonrival data. Giving data property rights to consumers can generate allocations that are close to optimal. Consumers balance their concerns for privacy against the economic gains that come from selling data broadly.
🚨Free data alert!! 🚨 Please share. Large new dataset of Amazon product reviews, including full text and photos and product characteristics, with individual *reviews labeled as fake reviews*. I believe this is the first publicly available data of this kind. github.com/bretthollenbeck/fake-reviews-…
No one wants to own their data until the platform where their data lives enshittifies to the point of unusability and/or starts selling their data in ways they didn't expect. Like, 99% of ppl are here because Elon owns Twitter. I get why people don't care, but they tend to eventually every time.