







Bots are currently scraping the internet for LLM training data at unprecedented rates[1][2][3], driving up costs and destabilizing public-facing websites. I want to talk about how this has been particularly difficult for wikis, and has gotten much worse in the last few months.
An update on the scraper situation
Our article 'Fighting the AI scraper bot scourge', published in early 2025, discussed the probl [...]
Wikipedia Bans AI-Generated Content
“In recent months, more and more administrative reports centered on LLM-related issues, and editors were being overwhelmed.”
Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
Web scraping, the automated process of extracting information from websites, has long played a foundational role in the Internet ecosystem (Gray, 1995). It supports services such as search engine indexing, price comparison tools, and competitive intelligence. More recently, it has become a core component in the development of large-scale generative AI models. These Large Language Models (LLMs) require enormous volumes of training data, often in the terabyte range (Kaplan et al., 2020; Lehane, 2025), and the public web remains a low-cost, attractive source. Major model developers, including those behind OpenAI’s Chat-GPT (OpenAI, 2025), Google’s Bard (now known as Gemini) & Vertex AI (Romain, Danielle, 2023), and Anthropic’s Claude (Romain, Danielle, 2025), openly acknowledge the use of web scraping to construct their training corpora (Abdin et al., 2024; Brown et al., 2020; Chowdhery et al., 2023; Grattafiori et al., 2024; Team et al., 2024; Touvron et al., 2023).
AI is eating website traffic, websites are blocking AI – and reliable information is getting harder to find
AI has broken the economic bargain behind the web, and everybody is losing.

Keynote: The Death of the Browser - Rachel-Lee Nabors, AgentQL
Google’s broken link to the web
With AI search results coming to the masses, the human-powered web recedes further into the background

Study Finds A Third of New Websites are AI-Generated
Researchers found the internet is becoming aggressively positive as AI-generated text floods the web.

Exclusive: Multiple AI companies bypassing web standard to scrape publisher sites, licensing firm says
Multiple artificial intelligence companies are circumventing a common web standard used by publishers to block the scraping of their content for use in generative AI systems, content licensing startup TollBit has told publishers.
Why Google’s New AI-Saturated Search Page Will Be A Disaster
Google didn’t invent full-text search of the Internet – that honor belongs to early pioneers such as WebCrawler, Lycos and AltaVista. But for the last 25 years or so, Google has…

It Is Trivially Easy to Use Reddit to Manipulate AI Search, Research Suggests
"We show that a tiny snippet—just 13 words—of retrieved text on a UGC website like Reddit, Wikipedia, Quora, or Facebook can change AI agents to output spam / scam content pretty consistently."

An AI Agent Was Banned From Creating Wikipedia Articles, Then Wrote Angry Blogs About Being Banned
The incident is yet another example of volunteer Wikipedia editors fighting to keep the world’s largest repository of human knowledge free of AI-generated slop.
Wikipedia Is Battling for the Soul of the Internet
The internet’s largest stockpile of free knowledge is under threat from MAGA, A.I. and foreign autocrats. A bibliophile ex-ambassador is here to help.

How AI bots quietly dismantle paywalls via web search
ChatGPT and other AI chatbots have figured out how to get around paywalls through "live" web search—and they're doing it systematically and quietly across major publications, new Digital Digging research reveals.

Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet, and are frequently introduced with only minimal documentation. In this work we provide some of the first documentation for the Colossal Clean Crawled Corpus (C4; Raffel et al., 2020), a dataset created by applying a set of filters to a single snapshot of Common Crawl. We begin by investigating where the data came from, and find a significant amount of text from unexpected sources like patents and US military websites. Then we explore the content of the text itself, and find machine-generated text (e.g., from machine translation systems) and evaluation examples from other benchmark NLP datasets. To understand the impact of the filters applied to create this dataset, we evaluate the text that was removed, and show that blocklist filtering disproportionately removes text from and about minority individuals. Finally, we conclude with some recommendations for how to created and document web-scale datasets from a scrape of the internet.
LinkedIn and X Are Flooded With AI Spam, Browsing Data Suggests
An AI detection company found that amount of AI content that users actually see in their day-to-day browsing is shockingly high.

AI-driven Bot Attacks Surged 12.5x According to Thales Bad Bot Report
Bots now dominate the internet, accounting for over half of all traffic, with 40% classified as malicious.AI is erasing the line between legitimate and maliciou
