







Publication-grade text justification for the web
the html review 05
the html review is an annual journal of literature made to exist on the web

2024-09-23-TPAC-Integrity-for-the-Web.pdf
Newspaper3k: Article scraping & curation — newspaper 0.0.2 documentation
Inspired by requests for its simplicity and powered by lxml for its speed:
The Consent Layer: Using ligatures to make web text expensive to scrape without asking
Publish for humans, not for crawlers. A web font that swaps the words in your HTML, so readers see your writing and AI training gets a stale copy.

Web 2.0 ... The Machine is Us/ing Us
No Longer No Sense of an Ending
This is my first feature piece on Contents Magazine, about harnessing the properties of hypertext and the Web for superior reader experiences and business results.
Long Live the Web: A Call for Continued Open Standards and Neutrality
The Web is critical not merely to the digital revolution but to our continued prosperity—and even our liberty. Like democracy itself, it needs defending

The Consent Layer: Using ligatures to make web text expensive to scrape without asking
ShieldFont is an open-source creative technology project that offers a practical opt-out from unauthorized AI training and disrupts what is collected when that choice is ignored. It swaps 45.8% of content words (around 24.4% of all words) in a page's source code for other (partially) random words, while the font restores the original text on screen. Readers see the work as intended; mass scrapers collect an altered version. In testing, shielding caused over 90% of pages that would otherwise pass the quality filter to be rejected, keeping them out of the training pipeline. Of those that still passed, 19.4% of all words conveyed false meaning, adding noise to unauthorized AI training datasets. This paper's goal is to walk newcomers through the whole process, in plain language and in order: the project's rationale, how it was built, the results, how to deploy it, and where to contribute.

Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet, and are frequently introduced with only minimal documentation. In this work we provide some of the first documentation for the Colossal Clean Crawled Corpus (C4; Raffel et al., 2020), a dataset created by applying a set of filters to a single snapshot of Common Crawl. We begin by investigating where the data came from, and find a significant amount of text from unexpected sources like patents and US military websites. Then we explore the content of the text itself, and find machine-generated text (e.g., from machine translation systems) and evaluation examples from other benchmark NLP datasets. To understand the impact of the filters applied to create this dataset, we evaluate the text that was removed, and show that blocklist filtering disproportionately removes text from and about minority individuals. Finally, we conclude with some recommendations for how to created and document web-scale datasets from a scrape of the internet.
FediForum | Growing the Open Social Web un-workshop, 2026/03/02

ShieldFont
Publish for humans, not for crawlers. A web font that swaps the words in your HTML, so readers see your writing and AI training gets a stale copy.

We wrote a @standard.site explainer — what it is & why it's useful 📝✨ ~shared standards for longform publishing? ~a way to see what friends read & write? ~the atmosphere's hottest collab? yes yes & yes! Seeing lots of q's after Bluesky's new standard.site link card integration. Here's a TL;DR!
What is Standard Site, and why is it useful?
lab.leaflet.pub