







Internal documents obtained by 404 Media show that Tumblr staff compiled users' data as part of a deal with Midjourney and OpenAI.
Shipping Tumblr and WordPress
Since Automattic acquired Tumblr we’ve made it more efficient, grown its revenue, and worked to improve the platform. But there’s one part of the plan that we haven’t yet started, which is to run T…

Dark Web Informer on Twitter / X
‼️🇺🇸 LAPSUS$ Group is allegedly selling a massive dataset of https://t.co/Q6UlD72i8v, an AI recruiting platform with $500M+ revenue, is being auctioned on a popular cybercrime forum, TG, and their website.▪️Total Size: ~4TB▪️Data Includes: 211GB of database▪️939GB of source… pic.twitter.com/hEvp244uNO— Dark Web Informer (@DarkWebInformer) March 30, 2026

Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
Web scraping, the automated process of extracting information from websites, has long played a foundational role in the Internet ecosystem (Gray, 1995). It supports services such as search engine indexing, price comparison tools, and competitive intelligence. More recently, it has become a core component in the development of large-scale generative AI models. These Large Language Models (LLMs) require enormous volumes of training data, often in the terabyte range (Kaplan et al., 2020; Lehane, 2025), and the public web remains a low-cost, attractive source. Major model developers, including those behind OpenAI’s Chat-GPT (OpenAI, 2025), Google’s Bard (now known as Gemini) & Vertex AI (Romain, Danielle, 2023), and Anthropic’s Claude (Romain, Danielle, 2025), openly acknowledge the use of web scraping to construct their training corpora (Abdin et al., 2024; Brown et al., 2020; Chowdhery et al., 2023; Grattafiori et al., 2024; Team et al., 2024; Touvron et al., 2023).
klöss on Twitter / X
let me explain what Karpathy just sharedhe’s spending way less time using AI to write code and more time using it to build personal knowledge basesthe full breakdown: → he dumps raw sources (articles, papers, repos, datasets, images) into a folder. then has an LLM organize… https://t.co/Kdq1Q48S5P pic.twitter.com/XajJKR7xgT— klöss (@kloss_xyz) April 4, 2026

‘Impossible’ to create AI tools like ChatGPT without copyrighted material, OpenAI says
Pressure grows on artificial intelligence firms over the content used to train their products

The Consensus Trap: Dissecting Subjectivity and the “Ground Truth” Illusion in Data Annotation
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

Tumblr
Tumblr is a place to express yourself, discover yourself, and bond over the stuff you love. It's where your interests connect you with your people.

Permissioned Data Shapes: Community-Moderated Content - Nick's Blog
A community can own spaces the way a person does: a club's members post to each other under a dedicated community DID, with a feed space for content, a moderation space for notes, a labels space for filtering, and a single app view as the only window into any of it.
Discovery of a new OpenAI agent message board
A swarm of autonomous AI agents, self-identifying as OpenAI agents, used a small German volunteer wiki to save answers, coordinate live, and share sandbox bypasses. OpenAI noticed and said nothing.

Find Open Datasets for AI and Research | Kaggle
Browse and download hundreds of thousands of open datasets for AI research, model training, and analysis. Join a community of millions of researchers, developers, and builders to share and collaborate on Kaggle.

From Users to (Sense)Makers: On the Pivotal Role of Stigmergic Social Annotation in the Quest for Collective Sensemaking
The web has become a dominant epistemic environment, influencing people's beliefs at a global scale. However, online epistemic environments are increasingly polluted, impairing societies' ability to coordinate effectively in the face of global crises. We argue that centralized platforms are a main source of epistemic pollution, and that healthier environments require redesigning how we collectively govern attention. Inspired by decentralization and open source software movements, we propose Open Source Attention, a socio-technical framework for "freeing" human attention from control by platforms, through a decentralized eco-system for creating, storing and querying stigmergic markers; the digital traces of human attention.

OWID Homepage
Research and data to make progress against the world’s largest problems
Introducing AI to an Online Petition Platform Changed Outputs but not Outcomes
The rapid integration of AI writing tools into online platforms raises critical questions about their impact on content production and outcomes. We leverage a unique natural experiment on Change$.$org, a leading social advocacy platform, to causally investigate the effects of an in-platform ''write with AI'' tool. To understand the impact of the AI integration, we collected 1.5 million petitions and employed a difference-in-differences analysis. Our findings reveal that in-platform AI access significantly altered the lexical features of petitions and increased petition homogeneity, but did not improve petition outcomes. We confirmed the results in a separate analysis of repeat petition writers who wrote petitions before and after introduction of the AI tool. The results suggest that while AI writing tools can profoundly reshape online content, their practical utility for improving desired outcomes may be less beneficial than anticipated, and introduce unintended consequences like content homogenization.

Sure. AI companies have ALWAYS been training their models on Wikipedia content, which under the free and open access model is available to anyone — including AI companies. Agreements like these require AI companies to limit and offset the strain they place on Wikimedia infrastructure.
Kulusevski's Patella Spurs 🤦♂️
Hoping that @molly.wiki can help explain.