







One of the most important AI copyright legal battles just took a major turn.
Firms like Meta and A16z admit having to pay billions for training data would ruin their generative-AI plans as they fight new copyright rules
Meta, Google, Microsoft, and Andreessen Horowitz are trying to keep AI developers from having to pay for copyrighted material used in AI training.
The Unbelievable Scale of AI’s Pirated-Books Problem
Meta pirated millions of books to train its AI. Search through them here.
Book publishers sue Meta over AI’s ‘word-for-word’ copying
Meta is accused of ripping copyrighted works from piracy websites.

Search LibGen, the Pirated-Books Database That Meta Used to Train AI
Millions of books and scientific papers are captured in the collection’s current iteration.
Meta made its own AI detection system. It should have just used Google’s
Content Seal has a lot of catching up to do.

Millions of Copyrighted Songs Were Fed to AI Music Generators – Now There's Proof
Millions of copyrighted songs trained AI music generators, and new searchable databases from The Atlantic now confirm which tracks were used.

Board Calls for New Rules on Deceptive AI During Conflicts
In analyzing the spread of AI-generated content in armed conflicts in a case on the 2025 Israel-Iran war, the Oversight Board calls on Meta to do more to allow users to identify such output.

Automated Justice: Issues, Benefits and Risks in the Use of Artificial Intelligence and Its Algorithms in Access to Justice and Law Enforcement
The use of artificial intelligenceArtificial Intelligence (AI) (AI) in the field of law has generated many hopes. Some have seen it as a way of relieving courts’ congestion, facilitating investigations, and making sentences for certain offences more consistent—and therefore fairer. But while it is true that the work of investigators and judges can be facilitated by these tools, particularly in terms of finding evidenceEvidence during the investigative process, or preparing legal summaries, the panorama of current uses is far from rosy, as it often clashes with the reality of field usage and raises serious questions regarding human rightsHuman rights. This chapter will use the RobodebtRobodebt Case to explore some of the problems with introducing automationAutomation into legal systems with little human oversight. AI—especially if it is poorly designed—has biases in its data and learning pathways which need to be corrected. The infrastructures that carry these tools may fail, introducing novel bias. All these elements are poorly understood by the legal world and can lead to misuse. In this context, there is a need to identify both the users of AIArtificial Intelligence (AI) in the area of law and the uses made of it, as well as a need for transparencyTransparency, the rules and contours of which have yet to be established.

Automated Justice: Issues, Benefits and Risks in the Use of Artificial Intelligence and Its Algorithms in Access to Justice and Law Enforcement
The use of artificial intelligenceArtificial Intelligence (AI) (AI) in the field of law has generated many hopes. Some have seen it as a way of relieving courts’ congestion, facilitating investigations, and making sentences for certain offences more consistent—and therefore fairer. But while it is true that the work of investigators and judges can be facilitated by these tools, particularly in terms of finding evidenceEvidence during the investigative process, or preparing legal summaries, the panorama of current uses is far from rosy, as it often clashes with the reality of field usage and raises serious questions regarding human rightsHuman rights. This chapter will use the RobodebtRobodebt Case to explore some of the problems with introducing automationAutomation into legal systems with little human oversight. AI—especially if it is poorly designed—has biases in its data and learning pathways which need to be corrected. The infrastructures that carry these tools may fail, introducing novel bias. All these elements are poorly understood by the legal world and can lead to misuse. In this context, there is a need to identify both the users of AIArtificial Intelligence (AI) in the area of law and the uses made of it, as well as a need for transparencyTransparency, the rules and contours of which have yet to be established.

Anthropic sued by authors over alleged misuse of copyrighted works for AI training
The complaint alleges that Anthropic used pirated versions of books by hundreds of thousands of authors to develop its AI models without proper authorization or compensation.

How a 40-Minute Window Brought Down a $10 Billion AI Startup: The Mercor Data Breach, Explained
A poisoned open-source package, a credential-stealing payload, and 4 terabytes of stolen data here’s what every AI company needs to learn…

Unlawful by design: Exposing the human rights costs of generative AI - Amnesty International
This briefing examines how standalone generative AI systems, based on unlawful web scraping, are in conflict with international human rights law (IHRL) and standards through their design, development and deployment. While these technologies promise sophisticated automation and efficiency, they rely on data collection and model training practices that abuse privacy rights, enable discrimination, and threaten […]

Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
Permissive licenses like MIT, Apache-2.0, and BSD-3-Clause dominate open-source AI, signaling that artifacts like models, datasets, and code can be freely used, modified, and redistributed. However, these licenses carry mandatory requirements: include the full license text, provide a copyright notice, and preserve upstream attribution, that remain unverified at scale. Failure to meet these conditions can place reuse outside the scope of the license, effectively leaving AI artifacts under default copyright for those uses and exposing downstream users to litigation. We call this phenomenon ``permissive washing'': labeling AI artifacts as free to use, while omitting the legal documentation required to make that label actionable. To assess how widespread permissive washing is in the AI supply chain, we empirically audit 124,278 dataset $\rightarrow$ model $\rightarrow$ application supply chains, spanning 3,338 datasets, 6,664 models, and 28,516 applications across Hugging Face and GitHub. We find that an astonishing 96.5\% of datasets and 95.8\% of models lack the required license text, only 2.3\% of datasets and 3.2\% of models satisfy both license text and copyright requirements, and even when upstream artifacts provide complete licensing evidence, attribution rarely propagates downstream: only 27.59\% of models preserve compliant dataset notices and only 5.75\% of applications preserve compliant model notices (with just 6.38\% preserving any linked upstream notice). Practitioners cannot assume permissive labels confer the rights they claim: license files and notices, not metadata, are the source of legal truth. To support future research, we release our full audit dataset and reproducible pipeline.

A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

The website that created an AI clone of its editor in chief
Every CEO Dan Shipper on doubling headcount while automating everything, building an agent out of 30,000 copyedits, and the “dirty secret” of writing with AI


People Are Crashing Out Over Sora 2’s New Guardrails

Bing Is Generating Images of SpongeBob Doing 9/11

OpenAI’s Sora 2 Copyright Infringement Machine Features Nazi SpongeBobs and Criminal Pikachus

OpenAI Can’t Fix Sora’s Copyright Infringement Problem Because It Was Built With Stolen Content