







Data dredging, also known as data snooping or p-hacking, is the misuse of data analysis to find patterns in data that can be presented as statistically significant, thus dramatically increasing and understating the risk of false positives. This is done by performing many statistical tests on the data and only reporting those that come back with significant results. Thus data dredging is also often a misused or misapplied form of data mining.
In the face of rampant AI, is ‘data poisoning’ a new form of civil disobedience?
Boycotts, sabotage and other types of civil disobedience have long served collective action against injustice.

In the face of rampant AI, is ‘data poisoning’ a new form of civil disobedience?
Boycotts, sabotage and other types of civil disobedience have long served collective action against injustice.

Monitoring Data Minimisation
Data minimisation is a privacy enhancing principle, stating that personal data collected should be no more than necessary for the specific purpose consented by the user. Checking that a program...

Data Feminism
Today, data science is a form of power. It has been used to expose injustice, improve health outcomes, and topple governments. But it has also been used to d...

The Data Minimization Principle in Machine Learning
The principle of data minimization aims to reduce the amount of data collected, processed or retained to minimize the potential for misuse, unauthorized access, or data breaches. Rooted in...

Defense Against Dishonest Charts
This is a guide to protect ourselves and to preserve what is good about turning data into visual things.

Data Supply Chains
Data is a critical resource. Like oil, gold, or lithium, both companies and countries covet data. Ultimately, like oil, data’s flow can enrich those that posses
Backdoor or Feature? A New Perspective on Data Poisoning
In a backdoor attack, an adversary adds maliciously constructed ("backdoor") examples into a training set to make the resulting model vulnerable to manipulation. Defending against such attacks---that is, finding and removing the backdoor examples---typically involves viewing these examples as outliers and using techniques from robust statistics to detect and remove them. In this work, we present a new perspective on backdoor attacks. We argue that without structural information on the training data distribution, backdoor attacks are indistinguishable from naturally-occuring features in the data (and thus impossible to ``detect'' in a general sense). To circumvent this impossibility, we assume that a backdoor attack corresponds to the strongest feature in the training data. Under this assumption---which we make formal---we develop a new framework for detecting backdoor attacks. Our framework naturally gives rise to a corresponding algorithm whose efficacy we show both theoretically and experimentally.
From Principle to Practice: Vertical Data Minimization for Machine Learning
Aiming to train and deploy predictive models, organizations collect large amounts of detailed client data, risking the exposure of private information in the event of a breach. To mitigate this,...

Microsoft “Digital Escorts” Could Expose Defense Dept. Data to Chinese Hackers — ProPublica
The Pentagon bans foreign citizens from accessing highly sensitive data, but Microsoft bypasses this by using engineers in China and elsewhere to remotely instruct American “escorts” who may lack expertise to identify malicious code.

DecryptAds — From Hidden Flows to Public Insight
DecryptAds maps programmatic ad-tech supply chains from public files — ads.txt, sellers.json, and related signals — so you can trace hidden flows and surface invalid supply.

AMD is investigating claims of stolen company data
The data for sale allegedly includes future products.

Data Leverage: A Framework for Empowering the Public in its...
Many powerful computing technologies rely on implicit and explicit data contributions from the public. This dependency suggests a potential source of leverage for the public in its relationship...

Research ethics: 3 ways to blow the whistle
Reporting suspicions of scientific fraud is rarely easy, but some paths are more effective than others.

I would agree that this Doctorow post not good. This paragraph in particular irks me. I think that a "database of earlier hacking challenges" implies plainly false things about what happened, like that the attack vector was not novel (a previously unknown-by-any-person vulnerability was exploited)
Will Stancil
No. This Cory Doctorow post is awful and borders on actual misinformation. The fact that a widely respected voice like him is spreading this kind of false narrative is, again, a giant blaring alarm about the effect the Bluesky info environment is having on people’s brains.