







For his Spurious Correlations project, Tyler Vigen compares data sets that are the very definition of “correlation is not causation”. For instance, the number o
Causality
Written by one of the preeminent researchers in the field, this book provides a comprehensive exposition of modern analysis of causation. It shows how causality has grown from a nebulous concept into a mathematical theory with significant applications in the fields of statistics, artificial intelligence, economics, philosophy, cognitive science, and the health and social sciences. Judea Pearl presents and unifies the probabilistic, manipulative, counterfactual, and structural approaches to causation and devises simple mathematical tools for studying the relationships between causal connections and statistical associations. Cited in more than 2,100 scientific publications, it continues to liberate scientists from the traditional molds of statistical thinking. In this revised edition, Judea Pearl elucidates thorny issues, answers readers' questions, and offers a panoramic view of recent advances in this field of research. Causality will be of interest to students and professionals in a wide variety of fields. Dr Judea Pearl has received the 2011 Rumelhart Prize for his leading research in Artificial Intelligence (AI) and systems from The Cognitive Science Society.

The C-Word: Scientific Euphemisms Do Not Improve Causal Inference From Observational Data
Causal inference is a core task of science. However, authors and editors often refrain from explicitly acknowledging the causal goal of research projects; they refer to causal effect estimates as associational estimates. This commentary argues that using the term “causal” is necessary to improve the quality of observational research. Specifically, being explicit about the causal objective of a study reduces ambiguity in the scientific question, errors in the data analysis, and excesses in the interpretation of the results.

Related work | Data Counterfactuals
Adjacent research areas and starting papers, curated in Semble and synced into the site.

Paper abstract: "Results showed that xx and yy were highly positively correlated." Researchers: what is your guess as to what correlation coefficient this translates to? (answer in next tweet)
Paper abstract: "Results showed that xx and yy were highly positively correlated."Researchers: what is your guess as to what correlation coefficient this translates to? (answer in next tweet)— Eiko Fried (@EikoFried) March 29, 2024
Viral Effects Are Not Network Effects
Virality and network effects are conflated by even experienced Founders, and it keeps them from developing the right strategies and playbooks.

Cosine similarity
In data analysis, cosine similarity is a measure of similarity between two non-zero vectors defined in an inner product space. Cosine similarity is the cosine of the angle between the vectors; that is, it is the dot product of the vectors divided by the product of their lengths. It follows that the cosine similarity does not depend on the magnitudes of the vectors, but only on their angle. The cosine similarity always belongs to the interval [ − 1 , + 1 ] . {\displaystyle [-1,+1].} For example, two proportional vectors have a cosine similarity of +1, two orthogonal vectors have a similarity of 0, and two opposite vectors have a similarity of −1. In some contexts, the component values of the vectors cannot be negative, in which case the cosine similarity is bounded in [ 0 , 1 ] {\displaystyle [0,1]} .
WRKSHP.tools | Cause-Effect Diagram
A cause-effect diagram is a visual tool designed to help you explore and identify possible causes for a problem or an observed phenomenon. This helps you to come up with possible root causes for the problem you observed, that you can then use to build experiments around.
Local Causation
The counterfactual and regularity theories are universal accounts of causation. I argue that these should be generalized to produce local accounts of causation. A hallmark of universal accounts of causation is the assumption that apparent variation in causation between locations must be explained by differences in background causal conditions, by features of the causal-nexus or causing-complex. The local account of causation presented here rejects this assumption, allowing for genuine variation in causation to be explained by differences in location. I argue that local accounts of causation are plausible, and have pragmatic, empirical and theoretical advantages over universal accounts. I then report on the use of presheaves as models of local causation. The use of presheaves as models of local variation has precedents in algebraic geometry, category theory and physics; they are here used as models of local causal variation. The paper presents this idea as stemming from an approach using presheaves as models of local truth. Finally, I argue that a proper balance between universal and local causation can be assuaged by moving from presheaves to fully-fledged sheaf models.
A Statistical Interrogation of “The Case for Causality, Part 1” by Rausch and Haidt – Matthew B. Jané
I don’t know anything about the literature on social media and mental health so my focus on this post is to interrogate the statistical approach taken by the article written by Zach Rausch and Jonathon Haidt (link here) and to some extent the original meta-analysis by Ferguson.


RCTs: Crucial Yet Crucially Limited
RCTs do NOT estimate average effects; they also rarely EXPLAIN that which they have demonstrated.

Is the Memory Shortage Intentional? | Contrary Research
A deep dive from Contrary Research.

[109] Data Falsificada (Part 1): "Clusterfake" - Data Colada
This is the introduction to a four-part series of posts detailing evidence of fraud in four academic papers co-authored by Harvard Business School Professor

This is why @atproto.science is so relevant rn This compilation of essays indicates that scientists are most frustrated by insufficient “community tools and resources... Essential infrastructure for sharing, maintaining and building on existing work and data is also badly underdeveloped." >
What Scientists Said: Results from Astera's First Essay Competition
asterainstitute.substack.comLet's look at one problem in Cook's paper; it's easy to see that the results are noise. Cook finds a negative effect of racial violence on patenting, using annual data. But when we use the state-year data to construct an annual dataset, the results disappear. 1/
Michael Wiebe
Should Lisa Cook resign from the Federal Reserve? No. Should her most famous paper be retracted? Yes.
This is crazy! The data from the study show patterns that make it nearly impossible for it to be legit. But we can learn a lot from the tone of the comments. "We have concerns about [x], please clarify." Out in the wild, this would be THESE IDIOTS FAKED THEIR STUDY!!
Health Nerd
This is one of the most remarkable academic debacles I've ever seen. A large RCT got published in BMJ. There are currently 44 Pubpeer comments, mostly about the data, including...well. Read for yourself. pubpeer.com/publications/C08779C45DB6E407…