







A growing set of web scrapers designed to output consistent geodata about as many places of business in the world as possible.
WDC - RDFa, Microdata, and Microformat Data Sets
More and more websites have started to embed structured data describing products, people, organizations, places, and events into their HTML pages using markup standards such as Microdata, JSON-LD, RDFa, and Microformats. The Web Data Commons project extracts this data from several billion web pages. So far the project provides 12 different data set releases extracted from the Common Crawls 2010 to 2023. The project provides the extracted data for download and publishes statistics about the deployment of the different formats.
Scraping Framework for Golang
Scraping framework for extracting the data you need from websites, used for a wide range of applications, like data mining, data processing or archiving

Common Crawl - Open Repository of Web Crawl Data
We build and maintain an open repository of web crawl data that can be accessed and analyzed by anyone.
Tim Berners-Lee, Giant Global Graph (2007)
A new vision for the web where the users master their data and link them to everybody's benefit in a giant global graph.
Website Informer
Website.informer: complete data lookup & free aggregated report on any domain including whois, visitors, IP & DNS details, competitors, owners, etc

Global Search Engine Market Share in the Top 15 GDP Nations (2026)
While Google takes a large share of the global search engine market, there are other search engines like Bing and Baidu that capture their share of international SEO.

Slushy - Products, Competitors, Financials, Employees, Headquarters Locations
Slushy provides tools for content creators in the digital content industry. Use the CB Insights Platform to explore Slushy's full profile.
OWID Homepage
Research and data to make progress against the worldโs largest problems
Cameron's World
๐ A web-collage of text and images excavated from the buried neighbourhoods of GeoCities.

The Missing Piece: How Location Data is Coming to the AT Protocol
A deep dive into community-driven geolocation schemas, emerging projects, and the future of location-aware social networking on Bluesky
Peak Web Has Passed | datagubbe.se
America's data center growth hot spots, mapped
Virginia and Texas lead the way, but they're becoming a major political fight.

This Tool Unmasks the Shadowy World of Ads that Track Your Location
It's usually very difficult to investigate the advertising industry. A new tool called DecryptAds aims to make it much easier with a massive dataset anyone can query.

The same-origin paradigm has unfortunately become a local maxima of the web. I'd argue the capture of the internet's search, social spaces, and data lock-in, as well as the lack of open protocols counter this capture can be directly attributed to building on top of the same-origin paradigm.
๐ฎ
Thinking about atproto as an identity and data *substrate* of the web. RSS couldn't go all the way because it built within the boundaries of the same-origin paradigm, just like the closed platforms that ended up enclosing most of the web.