







Read our beginner’s guide to attribution models to learn the difference between the different types and how to decide which one to use.
A Human-Centric Framework for Data Attribution in Large Language Models
In the current Large Language Model (LLM) ecosystem, creators have little agency over how their data is used, and LLM users may find themselves unknowingly plagiarizing existing sources. Attribution of LLM-generated text to LLM input data could help with these challenges, but so far we have more questions than answers: what elements of LLM outputs require attribution, what goals should it serve, how should it be implemented? We contribute a human-centric data attribution framework, which situates the attribution problem within the broader data economy. Specific use cases for attribution, such as creative writing assistance or fact-checking, can be specified via a set of parameters (including stakeholder objectives and implementation criteria). These criteria are up for negotiation by the relevant stakeholder groups: creators, LLM users, and their intermediaries (publishers, platforms, AI companies). The outcome of domain-specific negotiations can be implemented and tested for whether the stakeholder goals are achieved. The proposed approach provides a bridge between methodological NLP work on data attribution, governance work on policy interventions, and economic analysis of creator incentives for a sustainable equilibrium in the data economy.

MIRA 2026 - Modular Interoperable Research Attribution
Catalyzing Modular Interoperable Research Attribution - A workshop to design and prototype interoperable frameworks for modular research attribution. June 7-11, 2026 in Ireland.
MIRA 2026 - Modular Interoperable Research Attribution
Catalyzing Modular Interoperable Research Attribution - A workshop to design and prototype interoperable frameworks for modular research attribution. June 7-11, 2026 in Ireland.
Introduction to the Journal of Marketing Research Special Interdisciplinary Issue on Consumer Financial Decision Making
If you have citation software installed, you can download citation data to the citation manager of your choice

The Customer Discovery Handbook
This is a practioner's guide to conducting design research on personas, needfinding, value, and usability.
The AI Attribution Error
If an AI produces something useful it's because of your own skill in model choice, prompting, and steering. If not, it's because the model is a useless lying machine that can't follow directions.

WASA: WAtermark-based Source Attribution for Large Language Model-Generated Data
The impressive performances of Large Language Models (LLMs) and their immense potential for commercialization have given rise to serious concerns over the Intellectual Property (IP) of their training data. In particular, the synthetic texts generated by LLMs may infringe the IP of the data being used to train the LLMs. To this end, it is imperative to be able to perform source attribution by identifying the data provider who contributed to the generation of a synthetic text by an LLM. In this paper, we show that this problem can be tackled by watermarking, i.e., by enabling an LLM to generate synthetic texts with embedded watermarks that contain information about their source(s). We identify the key properties of such watermarking frameworks (e.g., source attribution accuracy, robustness against adversaries), and propose a source attribution framework that satisfies these key properties due to our algorithmic designs. Our framework enables an LLM to learn an accurate mapping from the generated texts to data providers, which sets the foundation for effective source attribution. Extensive empirical evaluations show that our framework achieves effective source attribution.
The Open Source Distributor Business Model
Dirk Riehle, dirk@riehle.org, Published in Computer vol. 54, no. 12 (December 2021), pp. 99-103. Abstract This article defines and discusses one particular commercial open source business model, ca…

CopyFair License - P2P Foundation
= a type of licensing or agreement that aims to re-introduce the principle and practice of reciprocity in markets that use mutualized knowledge (commons), by regulating contributions to these commons for those that commercialize it
A protocol for attributable modular research contribution - Matsuthoughts
Crafting a good (reasoning) model
A recent talk I gave on model training, reasoning, and the next frontier.

Content and Community
The old model for content sprung from geographic communities; the new model for content is to be the organizing principle for virtual communities.

Reader guide · Standard Reader
How to use Standard Reader — following, reading, listening, and making it yours. No technical knowledge needed.
leaflet reader should: - easily let you one-off longpost - aggregate bsky, rss, and standard.site post - let you take notes on everything, save anything to lists, feed into drafts - keep track of events u are rsvp'd to - let you subscribe to and participate in community/aggregators
From today, you can add attribution to your own saves in currents.is. Make sure to do it to give proper credit to the original author!
Part of the datacounterfactuals.org reading lists. Research on data provenance, dataset documentation, licensing and attribution audits, and technical source-attribution methods for understanding which data sources are available, permitted, or responsible for model behavior.
WASA: WAtermark-based Source Attribution for Large Language Model-Generated Data
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

Datasheets for Datasets
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

A large-scale audit of dataset licensing and attribution in AI