







‘Plain and aggregated search results such as URLs, snippets, and factual index data, are publicly accessible facts and are not "works protected under the Copyright Act." Google cannot use copyright law to block scraping of uncopyrighted search result data.’ seroundtable.com/google-lawsuit-serpapi-dismis…
Google Lawsuit Against SerpApi Over Scraping Search Results Has Been Dismissed
www.seroundtable.comJul 24, 2026 at 2:15 PM
Steal the Internet - Archiving Everything and Sharing It With Others
2014: 70% of the links within legal journals and 50% of the URLs from Supreme Court decisions did not contain the originally cited material.
Landmark German ruling declares Google's AI Overviews are Google's own words and makes it liable for false answers
A German regional court has ruled that Google is directly liable for the content of its AI search overviews. According to the court, previous limited liability protections for search engine operators don't apply to AI overviews. In this case, Google's AI had falsely linked two publishers to fraud and made claims that didn't appear in any of the linked sources. The ruling could set a precedent for AI-generated content liability worldwide.

A Former Google Engineer Built a Search Engine for Finding Every Privacy Violation You Face Online
Former Google engineer Tim Libert is releasing a search engine, webXray, that aims to find illicit online data collection and tracking—with the goal of becoming “the Henry Ford of tech lawsuits.”

Exclusive: Multiple AI companies bypassing web standard to scrape publisher sites, licensing firm says
Multiple artificial intelligence companies are circumventing a common web standard used by publishers to block the scraping of their content for use in generative AI systems, content licensing startup TollBit has told publishers.
Search privately and without ads — Uruky
Search privately and without ads using Uruky, the private search engine.

The Technical Feasibility of Divesting Google Chrome – Knight-Georgetown Institute
As the European Commission advances efforts under the Digital Markets Act to require Google to share its search data with competitors, lessons from historic antitrust remedies underscore how data access could be transformational in the AI-powered search market. While the Commission’s proposals represent a novel and comprehensive approach, key improvements to data scope and sharing frequency, privacy protections, and dispute resolution are needed. US courts and enforcers charged with implementing similar provisions should take note.

The Web Can Thrive Without Google’s Search Monopoly
Viable browsers and meaningful contributions to web standards can be sustained with more modest revenue streams, writes Alissa Cooper.

Common Crawl - Open Repository of Web Crawl Data
We build and maintain an open repository of web crawl data that can be accessed and analyzed by anyone.
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet, and are frequently introduced with only minimal documentation. In this work we provide some of the first documentation for the Colossal Clean Crawled Corpus (C4; Raffel et al., 2020), a dataset created by applying a set of filters to a single snapshot of Common Crawl. We begin by investigating where the data came from, and find a significant amount of text from unexpected sources like patents and US military websites. Then we explore the content of the text itself, and find machine-generated text (e.g., from machine translation systems) and evaluation examples from other benchmark NLP datasets. To understand the impact of the filters applied to create this dataset, we evaluate the text that was removed, and show that blocklist filtering disproportionately removes text from and about minority individuals. Finally, we conclude with some recommendations for how to created and document web-scale datasets from a scrape of the internet.
Google’s AI Overviews Can Scam You. Here’s How to Stay Safe
Beyond mistakes or nonsense, deliberately bad information being injected into AI search summaries is leading people down potentially harmful paths.

This Guy Has Built an Open Source Search Engine as an Alternative to Google in His Spare Time
"I found it very weird that there essentially is no way to browse the web in an open manner. So that's what I am trying to build," the founder of Stract said.

A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

Beyond APIs: Collecting Web Data for Research using the National Internet Observatory
Widespread Internet use offers unprecedented opportunities to study human behavior at scale, yet researchers face significant ethical and technical barriers when attempting to collect data for academic studies.
EFF to State AGs: Investigate Google's Broken Promise to Users
Google's Failure to Warn Users About Law Enforcement Demands for Data Is Deceptive

Finding real, quotable examples is key to my blog writing workflow. Because Google and the default web search produce a lot of junk, I use Exa, HackerNews, web search, what I've already written, and internal company docs to do this, and I recently added @semble.so to this stack.