







🚨Free data alert!! 🚨 Please share. Large new dataset of Amazon product reviews, including full text and photos and product characteristics, with individual *reviews labeled as fake reviews*. I believe this is the first publicly available data of this kind. github.com/bretthollenbeck/fake-reviews-…
Jul 11, 2025 at 9:18 PM

Fuji TV, Sankei Shimbun announce fabricated data found in their polls
TOKYO -- Fuji Television Network Inc. and the Sankei Shimbun Co. on June 19 announced the discovery that some data in 14 joint opinion polls they cond

Open Dataset | Yelp Data Licensing
The Yelp Open Dataset is a subset of Yelp data intended for educational use. It provides real-world business data, like reviews, photos, check-ins, and attributes.

Permissioned Data Diary 4: The Big Picture - Daniel's Leaflets
A special edition of the data diary that sketches out the rough shape of where we're heading.
Dark Web Informer on Twitter / X
‼️🇺🇸 LAPSUS$ Group is allegedly selling a massive dataset of https://t.co/Q6UlD72i8v, an AI recruiting platform with $500M+ revenue, is being auctioned on a popular cybercrime forum, TG, and their website.▪️Total Size: ~4TB▪️Data Includes: 211GB of database▪️939GB of source… pic.twitter.com/hEvp244uNO— Dark Web Informer (@DarkWebInformer) March 30, 2026

Living in Data: A Citizen's Guide to a Better Information Future (Paperback)
Jer Thorp’s analysis of the word “data” in 10,325 New York Times stories written between 1984 and 2018 shows a distinct trend: among the words most closely associated with “data,” we find not only its classic companions “information” and “digital,” but also a variety of new neighbors—from “scandal” and “misinformation” to “ethics,” “friends,” and “play.”To live in data in the twenty-first century

A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

Data Supply Chains
Data is a critical resource. Like oil, gold, or lithium, both companies and countries covet data. Ultimately, like oil, data’s flow can enrich those that posses
WDC - RDFa, Microdata, and Microformat Data Sets
More and more websites have started to embed structured data describing products, people, organizations, places, and events into their HTML pages using markup standards such as Microdata, JSON-LD, RDFa, and Microformats. The Web Data Commons project extracts this data from several billion web pages. So far the project provides 12 different data set releases extracted from the Common Crawls 2010 to 2023. The project provides the extracted data for download and publishes statistics about the deployment of the different formats.
Your Data, Your Control
How data portability can unlock competition and empower consumers January 15, 2026 Copyright and permission to reproduce For information on the Competition Bureau's activities, please contact: Information Centre Competition Bureau 50 Victoria Street Gatineau QC K1A 0C9
Democratizing Data
Democratizing Data builds a community-driven data ecosystem by identifying how datasets are used and reducing barriers to accessing high-quality public data. The initiative enhances the discoverability, usability, and relevance of data for researchers, policymakers, and stakeholder communities. A suite of tools and strategic partnerships supports this work by connecting users to the data, insights, and networks needed to inform decisions and generate impact.
Datasheets for Datasets
The machine learning community currently has no standardized process for documenting datasets, which can lead to severe consequences in high-stakes domains. To address this gap, we propose...

tfw @aaronstevenwhite.io brings an analysis as sharp as a knife to your half-baked Saturday-morning thoughts: aaronstevenwhite.leaflet.pub/3miwsz2hdv22i 🤯 If we're going to own our data, let's actually own our data. Which is to say: No, really, y'all, we're doing this. 💖🧠
Machine-readable attitudes - Computational Semantics++
aaronstevenwhite.leaflet.pubPREreview joins #LoveData26 celebration with a strong commitment to encouraging open peer review of diverse research outputs, including datasets. 📊 Try out our modular review workflow for datasets here: prereview.org/review-a-dataset Learn more: bit.ly/dataset-workflow @lovedataweek.bsky.social