







We need more signal, which means we want more noise. A lot of current scientific infrastructure is designed to minimize messiness: define a narrow question, collect the minimum data required to answer it, standardize the dataset, exclude complicating variables, finish the analysis, publish the result. That approach is understandable. It is also one reason we…
The signal and the noise: why so many predictions fail - but some don't
"Nate Silver's The Signal and the Noise is The Soul of …

The Limits of Data
Policymakers want to make decisions based on clear data, but important factors are lost when we rely solely on data. A philosopher writes:

Who Will Keep Research Data Infrastructure Open and Running?
The scientific community must consider the longevity of open research infrastructure—why it might fail and how to prevent it.

Who Will Keep Research Data Infrastructure Open and Running?
The scientific community must consider the longevity of open research infrastructure—why it might fail and how to prevent it.

Science Must Decentralize
Knowledge production doesn’t happen in a vacuum. Every great scientific breakthrough is built on prior work, and an ongoing exchange with peers in the field. That’s why we need to address the threat

The Engine of Scientific Discovery: How New Methods and Tools Spark Major Breakthroughs
Abstract. How do we spark new scientific discoveries? Why do some breakthroughs seem even accidental? And most importantly, how can we accelerate them and

The Data Minimization Principle in Machine Learning
The principle of data minimization aims to reduce the amount of data collected, processed or retained to minimize the potential for misuse, unauthorized access, or data breaches. Rooted in...

Rationality: What It Is, Why It Seems Scarce, Why It Matters by Steven Pinker
In the twenty-first century, humanity is reaching new heights of scientific understanding—and at ...
A Data Utopia for Science-of-Science
Here I want to briefly sketch out a vision for how to solve a key set of problems facing science-of-science researchers, using the relatively new idea of a ‘data trust.’ In my ideal wor…

Algorithmic Data Minimization for Machine Learning over...
Machine learning can analyze vast amounts of data generated by IoT devices to identify patterns, make predictions, and enable real-time decision-making. By processing sensor data, machine learning...

SoK: Data Minimization in Machine Learning
Data minimization (DM) describes the principle of collecting only the data strictly necessary for a given task. It is a foundational principle across major data protection regulations like GDPR...

Ultra-Processed Information: AI and the Coming Deluge of Noise | Frankly 128
Spatial Data Science
Data science is concerned with finding answers to questions on the basis of available data, and communicating that effort. Besides showing the results, this communication involves sharing the data used, but also exposing the path that led to the answers in a comprehensive and reproducible way. It also acknowledges the fact that available data may not be sufficient to answer questions, and that any answers are conditional on the data collection or sampling protocols employed.
How and When to Involve Crowds in Scientific Research | James Evans
The lone academic in a basement lab is an endangered species. Years ago I asked Paul Ginsparg, who founded arXiv, what features his machine-learning filter used to flag speculative submissions. The top three: single-authored, submitted on a weekend, and heavy citation of Newton and Einstein. Science is a contact sport now. Marion Poetz and Henry Sauermann's How and When to Involve Crowds in Scientific Research (https://lnkd.in/g5pwVbfD) is the field manual. What crowds can do: Volume. Zooniverse mobilizes 2.7 million people to classify galaxies and court records. Galaxy Zoo's co-founder hand-classified 50,000 galaxies in one week before concluding that isolation would break him. Reach. NASA spent years failing to predict solar flares, then broadcast the problem. The winner was a semi-retired radio engineer in rural New Hampshire who swapped satellite data for radio data. Experience. Patients and caregivers asked to generate research questions produced 826 that scientists had missed, including the link between aging and wound healing. Bench-to-bedside runs backward. One caution the book underplays: an open call is not an inclusive one. Participation costs in time, money, and access screen people out. Self-selection can narrow the very diversity that makes crowds thrive. As AI floods the labor supply of routine, homogeneous cognition, human crowds become more valuable, not less. Machines process. People notice what nobody thought to ask. Check out my review @ https://lnkd.in/gkzvWrRC and the book @ https://lnkd.in/g5pwVbfD!
This is why @atproto.science is so relevant rn This compilation of essays indicates that scientists are most frustrated by insufficient “community tools and resources... Essential infrastructure for sharing, maintaining and building on existing work and data is also badly underdeveloped." >
What Scientists Said: Results from Astera's First Essay Competition
asterainstitute.substack.comHypothesis: the simple act of each of us publicly sharing more what we're paying attention to (reading, watching, ..) would dramatically enhance collective sensemaking. In open source software they say "With enough eyes, all bugs are shallow"- perhaps there is a parallel for open source *attention*?