







A fast, catalog-first view of current initiatives shaping AI data access, licensing, and enforcement, with archived entries available on demand.
Datacurve | The data engine for frontier AI
Custom data for long-horizon reasoning, software engineering, and data science.

Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
Large Language Model (LLM) agents, acting on their users' behalf to manipulate and analyze data, are likely to become the dominant workload for data systems in the future. When working with data,...

New data agents across the Agentic Data Cloud | Google Cloud Blog
Learn about new data agents and tools for business analysts, data scientists, and database admins to integrate with the Agentic Data Cloud.

AI and Doctrinal Collapse
Artificial intelligence runs on data. But the two legal regimes that govern data—information privacy law and copyright law—are under pressure. Formally, each re
Databricks: Leading Data and AI Platform for Enterprises
Databricks offers a unified platform for data, analytics and AI. Build better AI with a data-centric approach. Simplify ETL, data warehousing, governance and AI on the Data Intelligence Platform.

A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

AI native industrial data platform for manufacturing | UMH
Standardize industrial data across sites and systems to reduce costs, improve efficiency and accelerate execution. Open-source, deployed at production sites across Europe, live in weeks.

Find Open Datasets for AI and Research | Kaggle
Browse and download hundreds of thousands of open datasets for AI research, model training, and analysis. Join a community of millions of researchers, developers, and builders to share and collaborate on Kaggle.

Data Streaming for AI: From Extractive Training to Sovereign Infrastructure DWeb Camp 2026
AI systems are consuming the world's content without compensating its creators. This session explores data streaming as a new paradigm — where content flows to AI in real time, with built-in rights management, usage tracking, and fair compensation — and asks what it would take to make this infrastructure decentralized, sovereign, and governed by the communities it serves.
GitHub - Responsible-Dataset-Sharing/easy-dataset-share: A CLI tool that helps AI researchers share datasets responsibly.
A CLI tool that helps AI researchers share datasets responsibly. - Responsible-Dataset-Sharing/easy-dataset-share
Perspective Chapter: Fit for Purpose? Creative Commons Licensing for Research Data in the Age of Artificial Intelligence
Licensing is an important component of the re-usability of research data, itself part of the FAIR principles: without clear, machine-readable licensing, datasets risk becoming technically...

Government to ease data consent rules for AI development | The Asahi Shimbun Asia & Japan Watch
To accelerate artificial intelligence development, the government plans to relax consent requirements for access to personal information while introducing tougher penalties for intentional misuse.

We Can Just Build Things
Build the tools your community needs — production-grade, privacy-respecting freedom tech, made with an AI agent and grounded in a verified, values-aligned catalog (Nostr, AT Protocol, and beyond).

AI Data Hive: Denver | Meetup
🐝 AI Data Hive: Denver 🏔️Unite with Denver's brightest minds in AI and Data!Are you passionate about the future of Artificial Intelligence and the power of data? Do you work as a Data Scientist, Data Engineer, Machine Learning Engineer, Analyst, or in any profession that touches the revolutionary

Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
Permissive licenses like MIT, Apache-2.0, and BSD-3-Clause dominate open-source AI, signaling that artifacts like models, datasets, and code can be freely used, modified, and redistributed. However, these licenses carry mandatory requirements: include the full license text, provide a copyright notice, and preserve upstream attribution, that remain unverified at scale. Failure to meet these conditions can place reuse outside the scope of the license, effectively leaving AI artifacts under default copyright for those uses and exposing downstream users to litigation. We call this phenomenon ``permissive washing'': labeling AI artifacts as free to use, while omitting the legal documentation required to make that label actionable. To assess how widespread permissive washing is in the AI supply chain, we empirically audit 124,278 dataset $\rightarrow$ model $\rightarrow$ application supply chains, spanning 3,338 datasets, 6,664 models, and 28,516 applications across Hugging Face and GitHub. We find that an astonishing 96.5\% of datasets and 95.8\% of models lack the required license text, only 2.3\% of datasets and 3.2\% of models satisfy both license text and copyright requirements, and even when upstream artifacts provide complete licensing evidence, attribution rarely propagates downstream: only 27.59\% of models preserve compliant dataset notices and only 5.75\% of applications preserve compliant model notices (with just 6.38\% preserving any linked upstream notice). Practitioners cannot assume permissive labels confer the rights they claim: license files and notices, not metadata, are the source of legal truth. To support future research, we release our full audit dataset and reproducible pipeline.

FAIRdata.ai — FAIR Data Assessment
Assess your research data's FAIRness. Automated pipeline using F-UJI + Claude AI. Free to use.
