







red dwarf now supports full text search using @mackuba.eu 's lycan https://tangled.org/@mackuba.eu/lycan
Introduction
A complete search engine and RAG pipeline in your browser, server or edge network with support for full-text, vector, and hybrid search in less than 2kb.

feynon on Twitter / X
how is Claude not RAG? As per my understanding from your writeup it uses structured tool calls and keyword based (I assume semantic text based) text search, instead of a regular vector database approach, isn’t the latter retrieval too?— feynon (@feynon_) September 15, 2025
Readwise on Twitter / X
🆕 We shipped an MCP server! You can new query your Readwise highlights inside of Claude, Cursor and more: pic.twitter.com/3HdIm4TRzJ— Readwise (@readwise) May 1, 2025
Kuba's Journal - Launching Lycan - a search tool for your likes
You know that feeling when you’re trying to find a post (skeet) that you’ve seen some time ago, and you can’t find it? It happens to me all the time. Let’s say someone from the Bluesky dev team wrote some piece of technical detail, or some ATProto developer from the community posted about their project - I know more or less what I’m looking for, but not enough to find it in the global search. But I know I probably gave it a like, because I like everything 😅
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet, and are frequently introduced with only minimal documentation. In this work we provide some of the first documentation for the Colossal Clean Crawled Corpus (C4; Raffel et al., 2020), a dataset created by applying a set of filters to a single snapshot of Common Crawl. We begin by investigating where the data came from, and find a significant amount of text from unexpected sources like patents and US military websites. Then we explore the content of the text itself, and find machine-generated text (e.g., from machine translation systems) and evaluation examples from other benchmark NLP datasets. To understand the impact of the filters applied to create this dataset, we evaluate the text that was removed, and show that blocklist filtering disproportionately removes text from and about minority individuals. Finally, we conclude with some recommendations for how to created and document web-scale datasets from a scrape of the internet.
The Beginner’s Guide to Text Embeddings & Techniques | deepset Blog
Text embeddings represent human language to computers, enabling tasks like semantic search. Here, we introduce sparse and dense vectors in a non-technical way.

onlytopbro's Porn Gifs | RedGIFs
Find out more about onlytopbro and explore their 391 porn GIFs and images. Browse the millions of other porn GIFs and images free on RedGIFs
Why Google’s New AI-Saturated Search Page Will Be A Disaster
Google didn’t invent full-text search of the Internet – that honor belongs to early pioneers such as WebCrawler, Lycos and AltaVista. But for the last 25 years or so, Google has…

Google for Developers Blog - News about Web, Mobile, AI and Cloud
Explore LangExtract: a Gemini-powered, open-source Python library for reliable, structured information extraction from unstructured text with precise source grounding.

👋 Jan on Twitter / X
Introducing Jan-v1: 4B model for web search, an open-source alternative to Perplexity Pro.In our evals, Jan v1 delivers 91% SimpleQA accuracy, slightly outperforming Perplexity Pro while running fully locally.Use cases:- Web search- Deep ResearchBuilt on the new version… pic.twitter.com/YApIShOAHI— 👋 Jan (@jandotai) August 12, 2025
SamuelLHuber/pi-fff
pi extension that replaces built-in find and grep with FFF-powered fuzzy file and content search
New demo version of atsearch.network is here 💃 quick recap: any AT Proto app can invent its own record type: scrobbles, books, photos, RPGs, etc... nothing searches across them so i built a search engine that reads your lexicon and figures out how to index you
New demo version of atsearch.network is here 💃 quick recap: any AT Proto app can invent its own record type: scrobbles, books, photos, RPGs, etc... nothing searches across them so i built a search engine that reads your lexicon and figures out how to index you
Is there a broadly used lexicon for *book* highlights yet? (eg. for your Xteink/KOReader, or highlighted.app, highlights) (I figure this could be handy for sense-making on @semble.so, and for everyone with their new fancy xteink devices…)