







Testing data quality effectively
Learn how focusing on user value and trust gives you a clearer, more effective way to test data quality
Data Integration
This book is an introduction to the problem of data integration and a rigorous account of one of the leading approaches to solving this problem.

MyData
The human-centric approach to data is aimed at a fair, sustainable, and prosperous digital society. In such a society, people get value from their data and set the agenda on how it is used. And for organisations, the ethical use of data is always the most attractive option.

Show me the data - Design Notes
Living in Data: A Citizen's Guide to a Better Information Future (Paperback)
Jer Thorp’s analysis of the word “data” in 10,325 New York Times stories written between 1984 and 2018 shows a distinct trend: among the words most closely associated with “data,” we find not only its classic companions “information” and “digital,” but also a variety of new neighbors—from “scandal” and “misinformation” to “ethics,” “friends,” and “play.”To live in data in the twenty-first century

Data Supply Chains
Data is a critical resource. Like oil, gold, or lithium, both companies and countries covet data. Ultimately, like oil, data’s flow can enrich those that posses
Designing Data-Intensive Applications (DDIA) — an O’Reilly book by Martin Kleppmann (The Wild Boar Book)
NoSQL… Big Data… Scalability… CAP Theorem… Eventual Consistency… Sharding…
FAIR Principles - GO FAIR
In 2016, the ‘FAIR Guiding Principles for scientific data management and stewardship’ were published in Scientific Data. The authors intended to provide guidelines to improve the Findability, Accessibility, Interoperability, and Reuse of digital assets. The principles emphasise machine-actionability (i.e., the capacity of… Continue reading →

The Limits of Data
Policymakers want to make decisions based on clear data, but important factors are lost when we rely solely on data. A philosopher writes:

Designing Data-Intensive Applications
Data is at the center of many challenges in system design today. Difficult issues need to be figured out, such as scalability, consistency, reliability, efficiency, and... - Selection from Designing Data-Intensive Applications [Book]
What makes something data?
This is a question I posted on BlueSky on Friday 11/21/25, inspired by a talk I recently attended about evaluation of “AI” systems. I think…

Linked Open Vocabularies (LOV)
Your entry point to high quality and reusable Vocabularies to describe Linked Data.


Assessment of the energy performance and sustainability of data centres in EU: first technical report.
This document represents the deliverable ‘First Technical Report’ for the project Study on Technical Assistance in support of implementing Article 12(5) of Directive 2023/1791 on the energy performance and sustainability of data centres, and has been produced by EY, AIT and Borderstep for the European Commission. The technical report assesses data centre energy efficiency and sustainability, and the reporting scheme of the Delegated Regulation 2024/1364. It uses reported data, supplemented where necessary, and evaluates energy performance via sustainability and key performance indicators. The report also addresses data quality and completeness and proposes improvements for the reporting scheme, including in terms of the reported information and indicators and user experience.
Datasets Guide | Unsloth Documentation
Learn how to create & prepare a dataset for fine-tuning.

Data Distribution Valuation
Data valuation is a class of techniques for quantitatively assessing the value of data for applications like pricing in data marketplaces. Existing data valuation methods define a value for a discrete dataset. However, in many use cases, users are interested in not only the value of the dataset, but that of the distribution from which the dataset was sampled. For example, consider a buyer trying to evaluate whether to purchase data from different vendors. The buyer may observe (and compare) only a small preview sample from each vendor, to decide which vendor's data distribution is most useful to the buyer and purchase. The core question is how should we compare the values of data distributions from their samples? Under a Huber characterization of the data heterogeneity across vendors, we propose a maximum mean discrepancy (MMD)-based valuation method which enables theoretically principled and actionable policies for comparing data distributions from samples. We empirically demonstrate that our method is sample-efficient and effective in identifying valuable data distributions against several existing baselines, on multiple real-world datasets (e.g., network intrusion detection, credit card fraud detection) and downstream applications (classification, regression).