







A CLI tool that helps AI researchers share datasets responsibly. - Responsible-Dataset-Sharing/easy-dataset-share
Find Open Datasets for AI and Research | Kaggle
Browse and download hundreds of thousands of open datasets for AI research, model training, and analysis. Join a community of millions of researchers, developers, and builders to share and collaborate on Kaggle.

Introducing FlexOlmo: a new paradigm for language model training and data collaboration | Ai2
Explore how FlexOlmo enables collaborative language model training without sacrificing data privacy or control, introducing a new, flexible approach to building shared AI models.
Can data collectives help strengthen vulnerable cultures in the face of AI?
"Data collectives and cooperatives, which let creators control the collection and distribution of their data, are emerging as preferred alternatives to big tech companies."
Unlocking Dataset Value for AI-Enabled Scientific Discovery (AI Datasets)
All proposals must be submitted in accordance with the requirements specified in the funding opportunity and in the Proposal & Award Policies & Procedures Guide (PAPPG) and its supplements.

Unlocking Dataset Value for AI-Enabled Scientific Discovery (AI Datasets)
All proposals must be submitted in accordance with the requirements specified in the funding opportunity and in the Proposal & Award Policies & Procedures Guide (PAPPG) and its supplements.

DataLicenses.org
Machine-readable hints for AI agents/crawlers; easy to adopt, rely on compliance.
Datacurve | The data engine for frontier AI
Custom data for long-horizon reasoning, software engineering, and data science.

A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

A Short Guide to Data Strikes and Conscious Data Contribution in the Context of 2026 Frontier AI
Back to the basics of data leverage.

Join | Mozilla Data Collective
Mozilla Data Collective is rebuilding the AI data ecosystem with communities at the centre.
Join | Mozilla Data Collective
Mozilla Data Collective is rebuilding the AI data ecosystem with communities at the centre.
Databricks: Leading Data and AI Platform for Enterprises
Databricks offers a unified platform for data, analytics and AI. Build better AI with a data-centric approach. Simplify ETL, data warehousing, governance and AI on the Data Intelligence Platform.

Datasheets for Datasets
The machine learning community currently has no standardized process for documenting datasets, which can lead to severe consequences in high-stakes domains. To address this gap, we propose...

Cline - AI Coding, Open Source and Uncompromised
Open-source AI coding agent with Plan/Act modes, MCP integration, and terminal-first workflows. Trusted by 5M+ developers worldwide.

We build AI that works for humans
Imbue builds AI to help people think, create, and build. We share our tools openly because we believe progress in AI should be collaborative and developer-driven

Got to talk at @aidotengineer.bsky.social conf last week about the need for collaborative AI engineering. All our current coding agents are single player. We're trying to scale up individual productivity, but creating tons of alignment problems in the process. We have no good tools for...