







Making Machine Learning a first-class web citizen
Introducing beginners to the mechanics of machine learning – Miriam Posner
Every year, I spend some time introducing students to the mechanics of machine learning with neural nets. I definitely don’t go into great depth; I usually only have one class for this. But I try to unpack at least some of the major concepts, so that ML isn’t quite such a black box.
ml5 - A friendly machine learning library for the web.
ml5.js aims to make machine learning approachable for a broad audience of artists, creative coders, and students. The library provides access to machine learning algorithms and models in the browser, building on top of TensorFlow.js with no other external dependencies.
The interplay between machine learning and data minimization under the GDPR: the case of Google’s topics API
The rapid speed of digitalization and the continuous technological disruption have lead to different paradoxes and contradictions. With the widespread adop

jax-js: an ML library for the web
JAX in pure JavaScript, as a flexible machine learning library and compiler.

Designing machine learning systems: an iterative process for production-ready applications
"Machine learning systems are both complex and unique. Complex because they consist of many different components and involve many different stakeholders. Unique because they're data dependent, with data varying wildly from one use case to the next. In this book, you'll learn a holistic approach to designing ML systems that are reliable, scalable, maintainable, and adaptive to changing environments and business requirements. Author Chip Huyen, co-founder of Claypot AI, considers each design decision--such as how to process and create training data, which features to use, how often to retrain models, and what to monitor--in the context of how it can help your system as a whole achieve its objectives. The iterative framework in this book uses actual case studies backed by ample references."--Amazon.com

Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
Web scraping, the automated process of extracting information from websites, has long played a foundational role in the Internet ecosystem (Gray, 1995). It supports services such as search engine indexing, price comparison tools, and competitive intelligence. More recently, it has become a core component in the development of large-scale generative AI models. These Large Language Models (LLMs) require enormous volumes of training data, often in the terabyte range (Kaplan et al., 2020; Lehane, 2025), and the public web remains a low-cost, attractive source. Major model developers, including those behind OpenAI’s Chat-GPT (OpenAI, 2025), Google’s Bard (now known as Gemini) & Vertex AI (Romain, Danielle, 2023), and Anthropic’s Claude (Romain, Danielle, 2025), openly acknowledge the use of web scraping to construct their training corpora (Abdin et al., 2024; Brown et al., 2020; Chowdhery et al., 2023; Grattafiori et al., 2024; Team et al., 2024; Touvron et al., 2023).
The Dark Forest and Generative AI
Proving you're a human on a web flooded with generative AI content

Personalized Machine Learning
This page contains collects information and supplementary material for my textbook Personalized Machine Learning:
Algorithmic Collective Action in Machine Learning
We initiate a principled study of algorithmic collective action on digital platforms that deploy machine learning algorithms. We propose a simple theoretical model of a collective interacting with a firm's learning algorithm. The collective pools the data of participating individuals and executes an algorithmic strategy by instructing participants how to modify their own data to achieve a collective goal. We investigate the consequences of this model in three fundamental learning-theoretic settings: the case of a nonparametric optimal learning algorithm, a parametric risk minimizer, and gradient-based optimization. In each setting, we come up with coordinated algorithmic strategies and characterize natural success criteria as a function of the collective's size. Complementing our theory, we conduct systematic experiments on a skill classification task involving tens of thousands of resumes from a gig platform for freelancers. Through more than two thousand model training runs of a BERT-like language model, we see a striking correspondence emerge between our empirical observations and the predictions made by our theory. Taken together, our theory and experiments broadly support the conclusion that algorithmic collectives of exceedingly small fractional size can exert significant control over a platform's learning algorithm.

Categories for Machine Learning
This seminar series seeks to promote the learning and use of Category Theory by Machine Learning Researchers

Playing with AI inference in Firefox Web extensions
The personal blog of Thomas Steiner
Ryan Cordell (@ryancordell.org)
Claude Code is blowing my mind—the 1st thing I’ve seen amid the AI hype that feels truly transformative One example—we’ve been doing genre classification work & I had the thought "it’d be nice to have a web application that lets users tag newspaper texts"—literal minutes later it exists & works
Google’s broken link to the web
With AI search results coming to the masses, the human-powered web recedes further into the background

Introducing Ringspace: A Proposal for the Human Web
For months, I've been working on a project to demonstrate how we can preserve humanity on the web. It's finally ready for testing.
