







Our article 'Fighting the AI scraper bot scourge', published in early 2025, discussed the probl [...]
Aggressive AI scrapers are making it kinda suck to run wikis
Bots are currently scraping the internet for LLM training data at unprecedented rates[1][2][3], driving up costs and destabilizing public-facing websites. I want to talk about how this has been particularly difficult for wikis, and has gotten much worse in the last few months.

Getting Bots to Respect Boundaries
Getting Bots to Respect Boundaries How AI Crawlers Are Straining Web Infrastructure Audrey HingleJanuary 2026 Image by Janet Turra & Cambridge Diversity Fundbetterimagesofai.org creativecommons.org/licenses/by/4.0 Contents Introduction: Setting up the Problem 3 Understandin...
How AI bots quietly dismantle paywalls via web search
ChatGPT and other AI chatbots have figured out how to get around paywalls through "live" web search—and they're doing it systematically and quietly across major publications, new Digital Digging research reveals.

Commence the Botwatch | Botwatch Blog
Bots were already a pain online when they were little more than if-then scripts. Now a tidal wave of slop is swamping the internet as big tech companies profit. OpenAI and the rest would love for you to believe they care about "alignment" and "values" while their AI models damage our communities, both online and off. It's time to fight back.

Keynote: The Death of the Browser - Rachel-Lee Nabors, AgentQL
AI-driven Bot Attacks Surged 12.5x According to Thales Bad Bot Report
Bots now dominate the internet, accounting for over half of all traffic, with 40% classified as malicious.AI is erasing the line between legitimate and maliciou

"Build protocols, not platforms"
Exploring the 'Protocols, Not Platforms' paper alongside a curated digest of recent AI crawler policy shifts and developer experiments.

Exclusive: Multiple AI companies bypassing web standard to scrape publisher sites, licensing firm says
Multiple artificial intelligence companies are circumventing a common web standard used by publishers to block the scraping of their content for use in generative AI systems, content licensing startup TollBit has told publishers.
Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
Web scraping, the automated process of extracting information from websites, has long played a foundational role in the Internet ecosystem (Gray, 1995). It supports services such as search engine indexing, price comparison tools, and competitive intelligence. More recently, it has become a core component in the development of large-scale generative AI models. These Large Language Models (LLMs) require enormous volumes of training data, often in the terabyte range (Kaplan et al., 2020; Lehane, 2025), and the public web remains a low-cost, attractive source. Major model developers, including those behind OpenAI’s Chat-GPT (OpenAI, 2025), Google’s Bard (now known as Gemini) & Vertex AI (Romain, Danielle, 2023), and Anthropic’s Claude (Romain, Danielle, 2025), openly acknowledge the use of web scraping to construct their training corpora (Abdin et al., 2024; Brown et al., 2020; Chowdhery et al., 2023; Grattafiori et al., 2024; Team et al., 2024; Touvron et al., 2023).
Patreon Blocks Crawlers From Stealing Creators' Work for AI Training
“Creators deserve credit, compensation, and consent. If that's not on the table, the crawlers can stay the fuck off Patreon," CEO Jack Conte wrote on Thursday.
Google’s broken link to the web
With AI search results coming to the masses, the human-powered web recedes further into the background

Katika Kühnreich (@Katika@chaos.social)
Attached: 1 image #SocialMedia of #realnameregistration , oh sorry, they call it real #humans , that advertises with " #AI " #slop #wsocial is more than problematic And already has #bots crawling it's #network Now they show that they are problematic concerning #aesthetics as well & advertise with tasteless #aislop @bildoperationen@tldr.nettime.org #realnameregistration is highly dangerous & if the company that wants to have your most private data can't save there net from bots they're not very trustworthy, are they? #tech

Stop Clicking on Junk: How to Identify and Ignore AI-Generated Slop Today
The internet is slowly being flooded by AI slop. What is it and how can you avoid it?

The website that created an AI clone of its editor in chief
Every CEO Dan Shipper on doubling headcount while automating everything, building an agent out of 30,000 copyedits, and the “dirty secret” of writing with AI

Zack Whittaker (@zackwhittaker@mastodon.social)
Another AI test gone awry, U.K. edition. "An agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code." More: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
Want to know more about the recent bot activity, what we've seen, and what we've done/are doing? Check out our write-up here! ✒️ 🔗 blog.tophhie.cloud/increased-bot-activity-what-w… Special mentions: @revfox84.bsky.social, @michelbestaat.thereforeiam.eu, @bmann.ca, and @baileytownsend.dev.
Increased Bot Activity & What We're Doing About It
blog.tophhie.cloud