







I’d like to raise a fun and interesting topic, and that’s intellectual property and content licensing. The data that flies around the Atmosphere is all visible and public. We’re at a stage within the Atmosphere where services are beginning to grow, and we’re beginning to see lexicons and data that express creative works that extend in length beyond microblogging. I think that the Standard.site lexicons are a great example. But these creations lead to further questions. Who has permission to...
A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

Licenses – Open Source Initiative
The content on this website, of which Opensource.org is the author, is licensed under a Creative Commons Attribution 4.0 International License.Opensource.org is not the author of any of the licenses reproduced on this site. Questions about the copyright in a license should be directed to the license steward. Read our Privacy Policy
Open Source Beyond Licensing - The Evolution Ahead
Boris Mann's digital garden and personal site.
your long-form atmosphere · pub search
what you publish, who reads it, your semantic neighborhoods, and what you recommend
A Human-Centric Framework for Data Attribution in Large Language Models
In the current Large Language Model (LLM) ecosystem, creators have little agency over how their data is used, and LLM users may find themselves unknowingly plagiarizing existing sources. Attribution of LLM-generated text to LLM input data could help with these challenges, but so far we have more questions than answers: what elements of LLM outputs require attribution, what goals should it serve, how should it be implemented? We contribute a human-centric data attribution framework, which situates the attribution problem within the broader data economy. Specific use cases for attribution, such as creative writing assistance or fact-checking, can be specified via a set of parameters (including stakeholder objectives and implementation criteria). These criteria are up for negotiation by the relevant stakeholder groups: creators, LLM users, and their intermediaries (publishers, platforms, AI companies). The outcome of domain-specific negotiations can be implemented and tested for whether the stakeholder goals are achieved. The proposed approach provides a bridge between methodological NLP work on data attribution, governance work on policy interventions, and economic analysis of creator incentives for a sustainable equilibrium in the data economy.


RSL: Really Simple Licensing
The open content licensing standard for the AI-first Internet
Perspective Chapter: Fit for Purpose? Creative Commons Licensing for Research Data in the Age of Artificial Intelligence
Licensing is an important component of the re-usability of research data, itself part of the FAIR principles: without clear, machine-readable licensing, datasets risk becoming technically...

Lexicon Embeds Overview
ATProto’s inside out structure provides an opportunity for open, permissionless building on top of any component in the network, including its extensible data schema definition system, Lexicons. One of the opportunities this presents is users sharing their data and records with multiple apps. Those apps may be direct substitutions (different clients for microblogging say), or apps that use that data in a variety of ways, such as Popfeed an app which aggregates posts about a variety of media ty...


Proposal: standard.site.declaration
Proposal: standard.site.declaration Draft for Discussion Abstract AT Protocol provides portable identity, portable content, and portable social graphs. What it currently lacks is a standard mechanism for discovering machine-readable declarations associated with digital content. Today, creators and organisations publish declarations about digital assets through websites, registries, the C2PA provenance framework, and other systems. These declarations may describe provenance, copyright ownership...


SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
The legality of training language models (LMs) on copyrighted or otherwise restricted data is under intense debate. However, as we show, model performance significantly degrades if trained only on low-risk text (e.g., out-of-copyright books or government documents), due to its limited size and domain coverage. We present SILO, a new language model that manages this risk-performance tradeoff during inference. SILO is built by (1) training a parametric LM on the Open License Corpus (OLC), a new corpus we curate with 228B tokens of public domain and permissively licensed text and (2) augmenting it with a more general and easily modifiable nonparametric datastore (e.g., containing copyrighted books or news) that is only queried during inference. The datastore allows use of high-risk data without training on it, supports sentence-level data attribution, and enables data producers to opt out from the model by removing content from the store. These capabilities can foster compliance with data-use regulations such as the fair use doctrine in the United States and the GDPR in the European Union. Our experiments show that the parametric LM struggles on its own with domains not covered by OLC. However, access to the datastore greatly improves out of domain performance, closing 90% of the performance gap with an LM trained on the Pile, a more diverse corpus with mostly high-risk text. We also analyze which nonparametric approach works best, where the remaining errors lie, and how performance scales with datastore size. Our results suggest that it is possible to build high quality language models while mitigating legal risk.
This approach to a data commons could extend to so many typically shared things. For businesses to easily and authoritatively publish things like open hours or product catalogs, discographies, every kind of -ography... creative appmakers could do so much with all that
JB Hutch
Testing an idea that I've had for a while, but I think is even better suited for atproto Internet, meet plts.at
Publishing services that lets a user publish content in the ATmosphere network. From long-form (blogging) to videos (vlogging), from images (galleries) to podcasting, and content. Not included are microblogging services, there's a separate list for it. #ATmosphereConf

TOKIMEKI Diary
WhiteWind atproto blog
Inkwell

AnyPub

Bulleted — an outliner for lists, notes, and plans
pixl