







Learn how focusing on user value and trust gives you a clearer, more effective way to test data quality
Overview data quality dimensions [Data Management Wiki]
A Data Utopia for Science-of-Science
Here I want to briefly sketch out a vision for how to solve a key set of problems facing science-of-science researchers, using the relatively new idea of a ‘data trust.’ In my ideal wor…

Identify trusted publishers for your research • Think. Check. Submit.
Tools and practical resources, aiming to educate researchers, promote integrity, and build trust in credible research and publications.

Establishing Trusted Representation | CAA
Labelling of quality and price is an important key for selecting goods or services. False or misleading representations may sway consumers into buying goods or services that are actually poor quality or overvalued. The Act against Unjustifiable Premiums and Misleading Representations prohibits such misleading representations. The Consumer Affairs Agency is addressing to ensure proper environment for shopping according to the Act.
MyData
The human-centric approach to data is aimed at a fair, sustainable, and prosperous digital society. In such a society, people get value from their data and set the agenda on how it is used. And for organisations, the ethical use of data is always the most attractive option.

FAIRdata.ai — FAIR Data Assessment
Assess your research data's FAIRness. Automated pipeline using F-UJI + Claude AI. Free to use.

Validation Free and Replication Robust Volume-based Data Valuation
Data valuation arises as a non-trivial challenge in real-world use cases such as collaborative machine learning, federated learning, trusted data sharing, data marketplaces. The value of data is often associated with the learning performance (e.g., validation accuracy) of a model trained on the data, which introduces a close coupling between data valuation and validation. However, a validation set may notbe available in practice and it can be challenging for the data providers to reach an agreement on the choice of the validation set. Another practical issue is that of data replication: Given the value of some data points, a dishonest data provider may replicate these data points to exploit the valuation for a larger reward/payment. We observe that the diversity of the data points is an inherent property of a dataset that is independent of validation. We formalize diversity via the volume of the data matrix (i.e., determinant of its left Gram), which allows us to establish a formal connection between the diversity of data and learning performance without requiring validation. Furthermore, we propose a robust volume measure with a theoretical guarantee on the replication robustness by following the intuition that copying the same data points does not increase the diversity of data. We perform extensive experiments to demonstrate its consistency in valuation and practical advantages over existing baselines and show that our method is model- and task-agnostic and can be flexibly adapted to handle various neural networks.
How to create software quality.
I’ve been reading Steven Sinofsky’s Hardcore Software, and particularly enjoyed this quote from a memo discussed in the Zero Defects chapter: You can improve the quality of your code, and if you do, the rewards for yourself and for Microsoft will be immense. The hardest part is to decide that you want to write perfect code. If I wrote that in an internal memo, I imagine the engineering team would mutiny, but software quality is certainly an interesting topic where I continue to refine my thinking. There are so many software quality playbooks out there, and I increasingly believe that all these playbooks work in their intended context, but are often misapplied.

Data Valuation in the Absence of a Reliable Validation Set
Data valuation plays a pivotal role in ensuring data quality and equitably compensating data contributors. Existing game-theoretic data valuation techniques mostly rely on the availability of a high-quality validation set for their efficacy. However, the feasibility of obtaining a clean validation set drawn from the test distribution may be limited in practice. In this work, we show that the choice of validation set can significantly impact the final data value scores. In order to mitigate this, we introduce a general paradigm that converts a traditional validation-based game-theoretic data valuation method into a validation-free alternative. Specifically, we utilize the cross-validation error as a surrogate for to evaluate the model's performance on a validation set. As computing the cross-validation error can be computationally expensive, we propose using the cross-validation error of a kernel regression model as an effective and efficient surrogate for the true performance score on the population. We compare the performance of the validation-free variant of existing data valuation techniques with their original validation-based counterparts. Our results indicate that the validation-free variants generally match or often significantly surpass the performance of their validation-based counterparts.
Establishing trust in automated reasoning - MetaROR
Since its beginnings in the 1940s, automated reasoning by computers has become a tool of ever growing importance in scientific research. So far, the rules underlying automated reasoning have mainly been formulated by humans, in the form of program source code. Rules derived from large amounts of data, via machine learning techniques, are a complementary approach currently under intense development. The question of why we should trust these systems, and the results obtained with their help, has been discussed by early practitioners of computational science, but was later forgotten. The present work focuses on independent reviewing, an important source of trust in science, and identifies the characteristics of automated reasoning systems that affect their reviewability. It also discusses possible steps towards increasing reviewability and trustworthiness via a combination of technical and social measures.

Data Distribution Valuation
Data valuation is a class of techniques for quantitatively assessing the value of data for applications like pricing in data marketplaces. Existing data valuation methods define a value for a discrete dataset. However, in many use cases, users are interested in not only the value of the dataset, but that of the distribution from which the dataset was sampled. For example, consider a buyer trying to evaluate whether to purchase data from different vendors. The buyer may observe (and compare) only a small preview sample from each vendor, to decide which vendor's data distribution is most useful to the buyer and purchase. The core question is how should we compare the values of data distributions from their samples? Under a Huber characterization of the data heterogeneity across vendors, we propose a maximum mean discrepancy (MMD)-based valuation method which enables theoretically principled and actionable policies for comparing data distributions from samples. We empirically demonstrate that our method is sample-efficient and effective in identifying valuable data distributions against several existing baselines, on multiple real-world datasets (e.g., network intrusion detection, credit card fraud detection) and downstream applications (classification, regression).
Don't Just Fact-Check Misinformation. First, Understand The Value It Holds. A Conversation w Researchers Michael Simeone & Kristy Roschke
Podcast Episode · Why Should I Trust You? · June 25 · 1h 7m
Scaling trust on the web
The Task Force for a Trustworthy Future Web's report on gaps and opportunities for how the next generation of online spaces will be built.

LLMs believe false statements even after explicit warnings that they're false
Fine-tuning tests show "bias... toward confidently representing the claims as true."

New study finds that when people help collect data or contribute to research it can build public trust by making scientists feel personally familiar and approachable, and that trust then spreads to how local and tangible the research feels. jcom.sissa.it/article/pubid/JCOM_2506_2026_…
How can citizen science reduce psychological distance to science? Insights from three projects in contested environmental contexts
jcom.sissa.it