







Evaluating how models perform across a range of knowledge work tasks, using live anonymized traffic and judged by models from Anthropic, OpenAI, and Google. An ongoing experiment.
Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
Web scraping, the automated process of extracting information from websites, has long played a foundational role in the Internet ecosystem (Gray, 1995). It supports services such as search engine indexing, price comparison tools, and competitive intelligence. More recently, it has become a core component in the development of large-scale generative AI models. These Large Language Models (LLMs) require enormous volumes of training data, often in the terabyte range (Kaplan et al., 2020; Lehane, 2025), and the public web remains a low-cost, attractive source. Major model developers, including those behind OpenAI’s Chat-GPT (OpenAI, 2025), Google’s Bard (now known as Gemini) & Vertex AI (Romain, Danielle, 2023), and Anthropic’s Claude (Romain, Danielle, 2025), openly acknowledge the use of web scraping to construct their training corpora (Abdin et al., 2024; Brown et al., 2020; Chowdhery et al., 2023; Grattafiori et al., 2024; Team et al., 2024; Touvron et al., 2023).
Infoenon — Where Information Meets Phenomenon
Search publisher-controlled information, public knowledge, and the open social web.

The Adoption and Usage of AI Agents: Early Evidence from Perplexity
This paper presents the first large-scale field study of the adoption, usage intensity, and use cases of general-purpose AI agents operating in open-world web environments. Our analysis centers on Comet, an AI-powered browser developed by Perplexity, and its integrated agent, Comet Assistant. Drawing on hundreds of millions of anonymized user interactions, we address three fundamental questions: Who is using AI agents? How intensively are they using them? And what are they using them for? Our findings reveal substantial heterogeneity in adoption and usage across user segments. Earlier adopters, users in countries with higher GDP per capita and educational attainment, and individuals working in digital or knowledge-intensive sectors -- such as digital technology, academia, finance, marketing, and entrepreneurship -- are more likely to adopt or actively use the agent. To systematically characterize the substance of agent usage, we introduce a hierarchical agentic taxonomy that organizes use cases across three levels: topic, subtopic, and task. The two largest topics, Productivity & Workflow and Learning & Research, account for 57% of all agentic queries, while the two largest subtopics, Courses and Shopping for Goods, make up 22%. The top 10 out of 90 tasks represent 55% of queries. Personal use constitutes 55% of queries, while professional and educational contexts comprise 30% and 16%, respectively. In the short term, use cases exhibit strong stickiness, but over time users tend to shift toward more cognitively oriented topics. The diffusion of increasingly capable AI agents carries important implications for researchers, businesses, policymakers, and educators, inviting new lines of inquiry into this rapidly emerging class of AI capabilities.

From Attention to Intention - The Open Garden
Exploring business models for atproto and the open social web
See what you think
Allegra A. Beal Cohen's blog about knowledge curation, new interfaces, and large-scale qualitative data.

The rise of community-curated knowledge networks
The Internet put thousands of years of human thought at our fingertips, and enabled billions of people to create content. At least 2.5…

The Politics of Open Infrastructures: Power, Governance, and Justice in Digital Knowledge Practices
This volume examines how openness is designed, governed, contested and lived in contemporary digital knowledge infrastructures. From open source software and internet standards, to citizen science platforms, public sector data systems and alternative computing practices, the book shows that infrastructures are never neutral technical backbones.

Curated retrieval versus open web search in public AI information...
Public institutions increasingly use large language models (LLMs) to answer citizens' questions, often pairing a curated knowledge base with live web search, yet whether the sources behind these...

How the Open Knowledge Format can improve data sharing | Google Cloud Blog
Learn how the Open Knowledge Format helps secure data sharing and improves collaboration across teams with standardized documentation.

How the Open Knowledge Format can improve data sharing | Google Cloud Blog
Learn how the Open Knowledge Format helps secure data sharing and improves collaboration across teams with standardized documentation.

How NASA is Using Graph Technology and LLMs to Build a People Knowledge Graph
Missed NASA’s People Graph webinar? Catch the recap and see how graph technology and AI are shaping the future of workforce intelligence.

The Paradox of Open
Numerous organisations and initiatives have been launched with a belief in openness and free knowledge. Their proponents placed their bets on the combined power of networked information services and new governance models for the production and sharing of content and data. We – as members of this broad movement – were among those who believed it possible to leverage this combination of power and opportunity to build a more democratic society, unleashing the power of the internet to create universal access to knowledge and culture. For us, such openness meant not only freedom, but also presented a path to justice and equality.
Avoiding Digital Productivity Traps - Cal Newport
Last week in this newsletter, I summarized some interesting results from a study that analyzed the behavior of 164,000 knowledge workers. It found that introducing ... Read more

From Social Network to Sense Making
Sure. AI companies have ALWAYS been training their models on Wikipedia content, which under the free and open access model is available to anyone — including AI companies. Agreements like these require AI companies to limit and offset the strain they place on Wikimedia infrastructure.
Kulusevski's Patella Spurs 🤦♂️
Hoping that @molly.wiki can help explain.
TiddlyWiki v5.4.0

Feeling Overwhelmed by Normal Browsers? - Horse Browser
Seams - Wisdom is made Together
Margin - Chrome Web Store

Explore — Semble