







The top of https://rslcollective.org/license says: Below are instructions for making your content available for training AI models through the collective licensing terms negotiated by the RSL Colle...
What is RSL? | RSL: Really Simple Licensing
The open content licensing standard for the AI-first Internet
Training AI Agents with RL | Unsloth Documentation
Learn how to train AI agents for real-world tasks using Reinforcement Learning (RL).

Building a Smarter AI Agent with Neural RAG - Will Bryk, Exa.ai
RSL: Really Simple Licensing
The open content licensing standard for the AI-first Internet
Unlawful by design: Exposing the human rights costs of generative AI - Amnesty International
This briefing examines how standalone generative AI systems, based on unlawful web scraping, are in conflict with international human rights law (IHRL) and standards through their design, development and deployment. While these technologies promise sophisticated automation and efficiency, they rely on data collection and model training practices that abuse privacy rights, enable discrimination, and threaten […]

Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
Web scraping, the automated process of extracting information from websites, has long played a foundational role in the Internet ecosystem (Gray, 1995). It supports services such as search engine indexing, price comparison tools, and competitive intelligence. More recently, it has become a core component in the development of large-scale generative AI models. These Large Language Models (LLMs) require enormous volumes of training data, often in the terabyte range (Kaplan et al., 2020; Lehane, 2025), and the public web remains a low-cost, attractive source. Major model developers, including those behind OpenAI’s Chat-GPT (OpenAI, 2025), Google’s Bard (now known as Gemini) & Vertex AI (Romain, Danielle, 2023), and Anthropic’s Claude (Romain, Danielle, 2025), openly acknowledge the use of web scraping to construct their training corpora (Abdin et al., 2024; Brown et al., 2020; Chowdhery et al., 2023; Grattafiori et al., 2024; Team et al., 2024; Touvron et al., 2023).
SkillMD.ai - Discover & Create Agent Skills
The open directory for agent skills. Discover, create, and share SKILL.md files for AI coding assistants.

Curated retrieval versus open web search in public AI information...
Public institutions increasingly use large language models (LLMs) to answer citizens' questions, often pairing a curated knowledge base with live web search, yet whether the sources behind these...

skills/skills/skill-creator/SKILL.md at main · anthropics/skills
Public repository for Agent Skills. Contribute to anthropics/skills development by creating an account on GitHub.
Chroma - open-source search infrastructure for AI
Open-source search infrastructure for AI

Karpathy shares 'LLM Knowledge Base' architecture that bypasses RAG with an evolving markdown library maintained by AI
Karpathy proposes something simpler and more loosely, messily elegant than the typical enterprise solution of a vector database and RAG pipeline.

Agentic Search for Dummies — Benjamin Anderson
A simple, effective baseline for building AI search agents.

DASL: RASL — Retrieval of Arbitrary Structures & Links
RASL is a URL scheme used to identify content-addressed DASL resources along with a simple HTTP-based retrieval method.

Did you know you can search for R packages by what they do, as opposed to just their names? Look! 👉🏼 rwarehouse.netlify.app @kylieainslie.bsky.social made The Warehouse for just this purpose. Send it to a new R programmer today ❤️ because it's the resource we all wish we'd had at some point #rstats
Sure. AI companies have ALWAYS been training their models on Wikipedia content, which under the free and open access model is available to anyone — including AI companies. Agreements like these require AI companies to limit and offset the strain they place on Wikimedia infrastructure.
Kulusevski's Patella Spurs 🤦♂️
Hoping that @molly.wiki can help explain.