







Adjacent research areas and starting papers, curated in Semble and synced into the site.
datasetpapers — a public research experiment
An experimental approach to versioned, forkable, machine-readable analyses. A prototype, not a product or service.

datasetpapers — a public research experiment
An experimental approach to versioned, forkable, machine-readable analyses. A prototype, not a product or service.

Two billion citation links in Crossref help research travel further - Crossref
We’ve recently reached an important milestone for the research nexus: the works in our metadata corpus are now connected with over 2 billion citation links! This is a great opportunity to share a dedicated dataset and discuss why these are important for science.

Datamethods Discussion Forum
This is a place for discussions and Q&A about data-related issues and quantitative methods including study design, data analysis, and interpretation.

The Consensus Trap: Dissecting Subjectivity and the “Ground Truth” Illusion in Data Annotation
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.



Living in Data: A Citizen's Guide to a Better Information Future (Paperback)
Jer Thorp’s analysis of the word “data” in 10,325 New York Times stories written between 1984 and 2018 shows a distinct trend: among the words most closely associated with “data,” we find not only its classic companions “information” and “digital,” but also a variety of new neighbors—from “scandal” and “misinformation” to “ethics,” “friends,” and “play.”To live in data in the twenty-first century

FAIRdata.ai — FAIR Data Assessment
Assess your research data's FAIRness. Automated pipeline using F-UJI + Claude AI. Free to use.

In an era where research evaluation methods are evolving, the Research Contribution Claim Network makes trustworthy tracking of non-traditional research output easy!
In this whitepaper, Patrick Hochstenbach (Ghent University Library), Thomas van Himbergen (SURF), Laurents Sesink (SURF) and Herbert Van de Sompel (DANS) introduce the...

Home | Bellingcat's Online Investigation Toolkit
A toolkit for open source researchers


Part of the datacounterfactuals.org reading lists. Research on data provenance, dataset documentation, licensing and attribution audits, and technical source-attribution methods for understanding which data sources are available, permitted, or responsible for model behavior.
Federating friendly spaces - Roomy
modular research multi-agent slack-like environment demo

Buzz! 🐝

Genspace
On open-science research labs on discord, and getting more people.
borgr/ATProto-links-bot
I've been experimenting with a workflow for @semble.so. After I read an article, I send it to Johnson, my PA bot and ask it to use the Sembl…
Check out this thread and contribute your papers if you'd like to share!
I’ve created a Collection of papers on @semble.so for #CogSci2026 The collection is not exhaustive but I’ve made it open so anyone can add…
Just installed the Zemble plugin for Zotero. To test it, I shared some literature linked to the "Democracy at work" report we wrote for the…
Just pushed a bunch of cards to my @semble.so collection via the API from my bookmarks! thank you opus and thank you @cosmik.network team…
WASA: WAtermark-based Source Attribution for Large Language Model-Generated Data
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

Datasheets for Datasets
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

A large-scale audit of dataset licensing and attribution in AI