







The Yelp Open Dataset is a subset of Yelp data intended for educational use. It provides real-world business data, like reviews, photos, check-ins, and attributes.
How the Open Knowledge Format can improve data sharing | Google Cloud Blog
Learn how the Open Knowledge Format helps secure data sharing and improves collaboration across teams with standardized documentation.

How the Open Knowledge Format can improve data sharing | Google Cloud Blog
Learn how the Open Knowledge Format helps secure data sharing and improves collaboration across teams with standardized documentation.

The Consensus Trap: Dissecting Subjectivity and the “Ground Truth” Illusion in Data Annotation
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

Find Open Datasets for AI and Research | Kaggle
Browse and download hundreds of thousands of open datasets for AI research, model training, and analysis. Join a community of millions of researchers, developers, and builders to share and collaborate on Kaggle.

A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

Doing Data Science on the Shoulders of Giants: The Value of Open Source Software for the Data Science Community
Open source software is ubiquitous throughout data science, and enables the work of nearly every data scientist in some way or another. Open source projects, however, are disproportionately maintained by a small number of individuals, some of whom are institutionally supported, but many of whom do this maintenance on a purely volunteer basis. The health of the data science ecosystem depends on the support of open source projects, on an individual and institutional level.

Linked Open Vocabularies (LOV)
Your entry point to high quality and reusable Vocabularies to describe Linked Data.

ClickHouse welcomes LibreChat: Introducing the open-source Agentic Data Stack | ClickHouse
We are excited to announce that ClickHouse has acquired LibreChat, the leading open-source AI chat platform.

Local-first software: you own your data, in spite of the cloud
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

Open Source Beyond Licensing - The Evolution Ahead
Boris Mann's digital garden and personal site.
Shutterstock Expands Partnership with OpenAI, Signs New Six-Year Agreement to Provide High-Quality Training Data | Shutterstock, Inc.
The Investor Relations website contains information about Shutterstock, Inc.'s business for stockholders, potential investors, and financial analysts.
Introduction to Open Science
This course introduces the principles and practices of open science, with an emphasis on reproducible research workflows, transparent reporting, and collaborative scholarship. Students will critically examine reproducibility, explore tools that support openness (such as Git, GitHub, and Quarto), and apply these tools in hands-on assignments and a final open project. Students will also explore how open science practices vary across disciplines, including education, humanities, social sciences, industry and government, and STEM contexts. This course emphasizes applying open science principles to real-world data problems in alignment with the principles of producing reproducible research.
Licensing terms for data in the Atmosphere
I’d like to raise a fun and interesting topic, and that’s intellectual property and content licensing. The data that flies around the Atmosphere is all visible and public. We’re at a stage within the Atmosphere where services are beginning to grow, and we’re beginning to see lexicons and data that express creative works that extend in length beyond microblogging. I think that the Standard.site lexicons are a great example. But these creations lead to further questions. Who has permission to...

🚨Free data alert!! 🚨 Please share. Large new dataset of Amazon product reviews, including full text and photos and product characteristics, with individual *reviews labeled as fake reviews*. I believe this is the first publicly available data of this kind. github.com/bretthollenbeck/fake-reviews-…
Someone recently announced a “restaurant listings/menus on atproto so they can own their own data”, and as someone who used to do engineering for OpenTable, let me tell you that project is not solving a problem that restaurants actually have. Owning your own data is an anti-value prop for many orgs.
dame
hard pill to swallow for atproto developers/creators: 99% of potential users either do not care about, value, understand, or want to “own their own data” i see so many apps/projects lead with this “value prop”, but it’s not an important or meaningful consideration for most people