







Meta, Google, Microsoft, and Andreessen Horowitz are trying to keep AI developers from having to pay for copyrighted material used in AI training.
Meta Secretly Trained Its AI on a Notorious Piracy Database, Newly Unredacted Court Docs Reveal
One of the most important AI copyright legal battles just took a major turn.

Pluralistic: Copyright won't solve creators' Generative AI problem (09 Feb 2023)
The media spectacle of generative AI (in which AI companies' breathless claims of their software's sorcerous powers are endlessly repeated) has understandably alarmed many creative workers, a group that's already traumatized by extractive abuse by media and tech companies.
‘Impossible’ to create AI tools like ChatGPT without copyrighted material, OpenAI says
Pressure grows on artificial intelligence firms over the content used to train their products

Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
Permissive licenses like MIT, Apache-2.0, and BSD-3-Clause dominate open-source AI, signaling that artifacts like models, datasets, and code can be freely used, modified, and redistributed. However, these licenses carry mandatory requirements: include the full license text, provide a copyright notice, and preserve upstream attribution, that remain unverified at scale. Failure to meet these conditions can place reuse outside the scope of the license, effectively leaving AI artifacts under default copyright for those uses and exposing downstream users to litigation. We call this phenomenon ``permissive washing'': labeling AI artifacts as free to use, while omitting the legal documentation required to make that label actionable. To assess how widespread permissive washing is in the AI supply chain, we empirically audit 124,278 dataset $\rightarrow$ model $\rightarrow$ application supply chains, spanning 3,338 datasets, 6,664 models, and 28,516 applications across Hugging Face and GitHub. We find that an astonishing 96.5\% of datasets and 95.8\% of models lack the required license text, only 2.3\% of datasets and 3.2\% of models satisfy both license text and copyright requirements, and even when upstream artifacts provide complete licensing evidence, attribution rarely propagates downstream: only 27.59\% of models preserve compliant dataset notices and only 5.75\% of applications preserve compliant model notices (with just 6.38\% preserving any linked upstream notice). Practitioners cannot assume permissive labels confer the rights they claim: license files and notices, not metadata, are the source of legal truth. To support future research, we release our full audit dataset and reproducible pipeline.

Can Agentic AI Coding Tools Finally End Copyright For Software While Re-Inventing Open Source?
Most of the discussions about the impact of the latest generative AI systems on copyright have centered on text, images and video. That’s no surprise, since writers, artists and film-makers feel ve…

Creative commons licenses and copyright may not stop academic work being used to train AI - Impact of Social Sciences
Considering the legal standing of creative commons licenses & copyright, Martin Eve suggests legal protections for academic work are unlikely to be forthcoming.

Who Pays for the Commons?
Three and a half years after the emergence of generative AI as a new technology paradigm, there is broad agreement that AI companies have extracted enorm...

Why so many game developers don't want to use generative AI
With credits ranging from Dispatch and Marvel Rivals to Uncharted and Dragon Age, over 30 devs share their thoughts on gen AI

Anthropic sued by authors over alleged misuse of copyrighted works for AI training
The complaint alleges that Anthropic used pirated versions of books by hundreds of thousands of authors to develop its AI models without proper authorization or compensation.

A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

Unlawful by design: Exposing the human rights costs of generative AI - Amnesty International
This briefing examines how standalone generative AI systems, based on unlawful web scraping, are in conflict with international human rights law (IHRL) and standards through their design, development and deployment. While these technologies promise sophisticated automation and efficiency, they rely on data collection and model training practices that abuse privacy rights, enable discrimination, and threaten […]

An Elegant Solution to AI Slop: Tax It, and Use the Resulting Billions of Dollars to Fund Cultural Institutions, Artists, and Researchers
A technologist proposes slapping a "slop tax" on companies that use or create generative AI content.

Companies Are Throttling Employees’ AI Use Because It’s Too Expensive
Sources and leaks from Amazon, Adobe, Atlassian, Citi, and more show what is really happening with AI right now: companies are trying to rein in AI use as costs spiral out of control.

Nobody Wants to Pay for Your AI
AI is everywhere today. It writes social posts, suggests email replies, recommends your next binge-watch, and quietly powers tools you use without even noticing.
We're Not Building AI Features for the Money
From the Zed Blog: Why Zed invests in AI, and the future we're building toward.
AI Companies Are Trying to Hide a Staggering Amount of Debt
AI companies are pouring tens of billions of dollars into enormous data centers. They're being built on top of a mountain of hidden debt.
