







Finding real, quotable examples is key to my blog writing workflow. Because Google and the default web search produce a lot of junk, I use Exa, HackerNews, web search, what I've already written, and internal company docs to do this, and I recently added @semble.so to this stack.
Jul 8, 2026 at 12:22 PM
Google’s AI Overviews Can Scam You. Here’s How to Stay Safe
Beyond mistakes or nonsense, deliberately bad information being injected into AI search summaries is leading people down potentially harmful paths.

Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet, and are frequently introduced with only minimal documentation. In this work we provide some of the first documentation for the Colossal Clean Crawled Corpus (C4; Raffel et al., 2020), a dataset created by applying a set of filters to a single snapshot of Common Crawl. We begin by investigating where the data came from, and find a significant amount of text from unexpected sources like patents and US military websites. Then we explore the content of the text itself, and find machine-generated text (e.g., from machine translation systems) and evaluation examples from other benchmark NLP datasets. To understand the impact of the filters applied to create this dataset, we evaluate the text that was removed, and show that blocklist filtering disproportionately removes text from and about minority individuals. Finally, we conclude with some recommendations for how to created and document web-scale datasets from a scrape of the internet.

Google Search's guidance about AI-generated content | Google Search Central Blog | Google for Developers
In this post, we'll share more about how AI-generated content fits into our long-standing approach to show helpful content to people on Search.

You Should Still Fact-Check the 'Expert Advice' in Google's AI Summaries
Google's AI responses will now show expert advice pulled from online forums like Reddit with specific quotes and links to discussions related to search queries.

It Is Trivially Easy to Use Reddit to Manipulate AI Search, Research Suggests
"We show that a tiny snippet—just 13 words—of retrieved text on a UGC website like Reddit, Wikipedia, Quora, or Facebook can change AI agents to output spam / scam content pretty consistently."

Quotation errors in general science journals
Due to the incremental nature of scientific discovery, scientific writing requires extensive referencing to the writings of others. The accuracy of this referencing is vital, yet errors do occur. These errors are called ‘quotation errors’. This paper presents the first assessment of quotation errors in high-impact general science journals. A total of 250 random citations were examined. The propositions being cited were compared with the referenced materials to verify whether the propositions could be substantiated by those materials. The study found a total error rate of 25%. This result tracks well with error rates found in similar studies in other academic fields. Additionally, several suggestions are offered that may help to decrease these errors and make similar studies more feasible in the future.

our system of knowledge production is under attack from an extraordinary array of forces and if the people putatively on the side of maintaining it cannot commit to the most basic standards (not using fake citations, writing your own words, etc) I fear there is no hope
Dr. Holly Walters
I'm sorry, what? In writing my first monograph, I spent six weeks trying to track down a citation in TWO languages I didn't know. And good thing too, because the citation was wrong. That's scholarship. That's research. You know, the thing we're trained to do?!?
created custom letta tools so my agent go research and highlight, bookmark, annotate, and organize stuff into collections atproto is awesome
funferall
@margin.at is so cool
Yup
selfhosting.sh
Interesting signal: Tailscale's co-founder (apenwarr) confirmed Tailscale SSH should work with Headscale. If you're self-hosting Headscale as your coordination server, that means you can get Tailscale's SSH features without depending on their cloud at all. Full self-hosted mesh + SSH tunneling.
to give an update on where that proposal stands: I got as far as publishing demo lexicons and deploying an example that passes through headers aligned with the IETF AIPREF work last summer: demo.user-intents.org
Rude1 Haunted Badness. ⁂
Suppose a Bluesky user does not want any of their public data to be used for generative AI training. They would go in to app settings, find the data reuse preferences section, and configure “Generative AI” 🤖 to “disallow” 🚫. The app would create this public record in their repository:
Here's how we leverage @semble.so's wonderful service in search results 👀. This is currently the only working search integration but 20+ more are in the works right now!

The Age of PageRank is Over

Informatics of the Oppressed
Peer Knowledge Assisted Search Using Community Search Logs

A survey of community search over big graphs

Remembering the pre-Google web, when search was an experiment
Curated vs open search (by Ronen Tamari) — Semble
あ (@aiueo.ooo)

Google’s AI search is so broken it can ‘disregard’ what you’re looking for

Google Search’s AI evolution includes more ads

Searching for 'Disregard' Breaks Google [Updated]
あ (@aiueo.ooo)