







When I read lively prose, one thing I notice is that it packs a lot of information per sentence—but not, usually, by being compact. Instead, how to explain it?Take this phrase “an interesting and exciting new finding”—that’s pretty dead. And one of several reasons is that the…— Henrik Karlsson (@phokarlsson) February 24, 2026
How Compaction Works in Pi | EARENDIL
Why compaction is needed for large language models and how Pi implements it

Shane Littrell, PhD on Twitter / X
One of the red flags of AI-written text is the heavy use of these odd linguistic constructions called “cataphoric teasers” to create an artificial feeling of suspense. These are easy to spot because they usually take the form of phrases like, “Here’s the part that nobody tells… https://t.co/kT4jRUh13x— Shane Littrell, PhD (@MetacogniShane) August 26, 2026

Amelia Wattenberger 🪷 on Twitter / X
exploring ways we can "step back" from text/code and make sense of it,the way do in maps or how we can "fit more" in our visual system, it just gets smaller and decreases resolution.here's a fun one: extract the main concepts and throw them on a sphere, keep spinning to go… pic.twitter.com/yzh42EcBTe— Amelia Wattenberger 🪷 (@Wattenberger) July 6, 2026

Description not evaluation - The Cynefin Co
One of the things I am enjoying at the moment is the way a lot of things are coming together conceptually. It happens like this when you develop or repurpose knowledge from different sources. Individual practices and ideas make sense in their own right. Then you read more, practice more, and the various origins […]

Experimental evidence of the effects of large language models versus web search on depth of learning
Abstract. The effects of using large language models (LLMs) versus traditional web search on depth of learning are explored. A theory is proposed that when

The Continual Learning Problem
A perspective on continual learning, motivating our paper on sparse memory finetuning

The Cloister Web: Reshaping the Political Maidan
The advent of the "Cloister Web," a conceptual space where individuals leverage Large Language Models (LLMs) to cultivate novel ideas and commit them to a persistent public memory, heralds a profound shift in our intellectual and political landscapes.

Inducing Sustained Creativity and Diversity in Large Language Models
We address a not-widely-recognized subset of exploratory search, where a user sets out on a typically long "search quest" for the perfect wedding dress, overlooked research topic, killer company idea, etc. The first few outputs of current large language models (LLMs) may be helpful but only as a start, since the quest requires learning the search space and evaluating many diverse and creative alternatives along the way. Although LLMs encode an impressive fraction of the world's knowledge, common decoding methods are narrowly optimized for prompts with correct answers and thus return mostly homogeneous and conventional results. Other approaches, including those designed to increase diversity across a small set of answers, start to repeat themselves long before search quest users learn enough to make final choices, or offer a uniform type of "creativity" to every user asking similar questions. We develop a novel, easy-to-implement decoding scheme that induces sustained creativity and diversity in LLMs, producing as many conceptually unique results as desired, even without access to the inner workings of an LLM's vector space. The algorithm unlocks an LLM's vast knowledge, both orthodox and heterodox, well beyond modal decoding paths. With this approach, search quest users can more quickly explore the search space and find satisfying answers.

Everything is miscellaneous : the power of the new digital disorder
Includes bibliographical references (p. [235]-257) and index; Prologue : information in space -- The new order of order -- Alphabetization and its discontents -- The geography of knowledge -- Lumps and splits -- The laws of the jungle -- Smart leaves -- Social knowing -- What nothing says -- Messiness as a virtue -- The work of knowledge -- Coda : misc; Philosopher Weinberger shows how the digital revolution is radically changing the way we make sense of our lives. Human beings constantly collect, label, and organize data--but today, the shift from the physical to the digital is mixing, burning, and ripping our lives apart. In the past, everything had its one place--the physical world demanded it--but now everything has its places: multiple categories, multiple shelves. Everything is suddenly miscellaneous. Weinberger charts the new principles of digital order that are remaking business, education, politics, science, and culture. He examines how Rand McNally decides what information not to include in a physical map (and why Google Earth is winning that battle), how Staples stores emulate online shopping to increase sales, why your children's teachers will stop having them memorize facts, and how the shift to digital music stands as the model for the future.--From publisher description; From A to Z, Everything Is Miscellaneous will completely reshape the way you think - and what you know - about the world. Includes information on alphabetical order, Amaxon.com, animals, Aristotle, authority, Bettmann Archive, blogs (weblogs), books, broadcasting, British Broadcasting Corporation (BBC), business, card catalog, categories and categorization, clusters, companies, Colon Classification, conversation, Melvil Dewey, Dewey Decimal Classification system, Encyclopaedia Britannica, encyclopedia, essentialism, experts, faceted classification system, first order of order, Flickr.com, Google, Great Books of the Western World, ancient Greeks, health and medical information, identifiers, index, inventory tracking, knowledge, labels, leaf and leaves, libraries, Library of Congress, links, Carolus Linnaeus, lumping and splitting, maps and mapping, marketing, meaning, metadata, multiple listing services (MLS), names of people, neutrality or neutral point of view, New York Public Library, Online Computer Library Center (OCLC), order and organization, people, physical space, everything having place, Plato, race, S.R. Ranganathan, Eleanor Rosch, Joshua Schacter, science, second order of order, simplicity, social constructivism, social knowledge, social networks, sorting, species, standardization, tags, taxonomies, third order of roder, topical categorization, tree, Uniform Product Code (UPC), users, Jimmy Wales, web, Wikipedia, etc

Experimental evidence of the effects of large language models versus web search on depth of learning
Abstract The effects of using large language models (LLMs) versus traditional web search on depth of learning are explored. A theory is proposed that when individuals learn about a topic from LLM syntheses, they risk developing shallower knowledge than when they learn through standard web search, even when the core facts in the results are the same. This shallower knowledge accrues from an inherent feature of LLMs—the presentation of results as summaries of vast arrays of information rather than individual search links—which inhibits users from actively discovering and synthesizing information sources themselves, as in traditional web search. Thus, when subsequently forming advice on the topic based on their search, those who learn from LLM syntheses (vs. traditional web links) feel less invested in forming their advice, and, more importantly, create advice that is sparser, less original, and ultimately less likely to be adopted by recipients. Results from seven online and laboratory experiments (n = 10,462) lend support for these predictions, and confirm, for example, that participants reported developing shallower knowledge from LLM summaries even when the results were augmented by real-time web links. Implications of the findings for recent research on the benefits and risks of LLMs, as well as limitations of the work, are discussed.

Human Intelligence, the Secret of Artificial Intelligence
Artificial intelligence is mysterious: we speak to it and it seems to understand what we say. Proof that it understands is that it responds with text or speech that makes sense, and sometimes more …

Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet, and are frequently introduced with only minimal documentation. In this work we provide some of the first documentation for the Colossal Clean Crawled Corpus (C4; Raffel et al., 2020), a dataset created by applying a set of filters to a single snapshot of Common Crawl. We begin by investigating where the data came from, and find a significant amount of text from unexpected sources like patents and US military websites. Then we explore the content of the text itself, and find machine-generated text (e.g., from machine translation systems) and evaluation examples from other benchmark NLP datasets. To understand the impact of the filters applied to create this dataset, we evaluate the text that was removed, and show that blocklist filtering disproportionately removes text from and about minority individuals. Finally, we conclude with some recommendations for how to created and document web-scale datasets from a scrape of the internet.
Ever thought we acquire generalizable knowledge by discarding details and compressing our experiences? In a new BBS paper, @sabinasloman.bsky.social and I argue otherwise, proposing a novel way of studying human learning inspired by double descent in ML. Disagree? Propose a commentary by May 15 :)
For me personally there is a journey here - I was a firm believer in archive everything. This peaked in a period where I was very interested in IPFS, CIDs, CAS, etc. Now, the next step. Post-Archive(-Everything). Through a series of conversations I came to the conclusion that forgetting is probably more powerful than remembering. The links in this collections play a similar note, or are related to my understanding of this.

FOSDEM 2026 - Willow - Protocols for an uncertain future
Willow - Home

I Deleted My Second Brain