







Conservation of information sparked scientific interest once a recurring pattern was noticed in the evolutionary computing literature. In grappling with the creation of information through evolutionary algorithms, this literature consistently revealed that the information outputted by such algorithms always needed first to be programmed into them. Thus, the primary goal of this literature—to uncover how information could be created from scratch or de novo —was shown to be misconceived: the information was not created but instead shuffled around or smuggled in, implying that it already existed in some form or other. Information output in these situations therefore always presupposed a counterbalancing input of prior information. Once this pattern was seen, the next logical step was to quantify the amount of information inputted and outputted, demonstrating a consistent mathematical relation between the two. This led to the proof of a number of theorems about search. In these theorems, a baseline search with probability p of success gave way to an improved search with probability q of success. Typically p would be very small and close to zero, implying a practically impossible search (like searching for a needle in a haystack). By contrast, q would be much larger and close to one, implying an eminently doable search. The punchline of these theorems was that, as the improved search became itself the subject of a search (a search for a search , or S4S), the probability of finding it could not exceed p / q , rendering success of the improved search no more probable than success of the original baseline search, in effect filling one hole by digging another. Such conservation-of-information theorems, as they came to be called, were search-space specific, adapted to different kinds of search across a range of search spaces. There was a measure-theoretic theorem in which probability measures guided search. There were also function-theoretic and fitness-theoretic theorems where mappings into the search space as well as fitness functions on the search space respectively guided search. The key insight of this paper is that all these conservation-of-information theorems are special cases of a simple probabilistic relation based on elementary probability theory. This paper identifies the underlying rationale that makes all the previous conservation-of-information theorems work. In so doing, it provides a straightforward proof and general formulation of what may rightly be called the Law of Conservation of Information.
Genetic Information: A Metaphor In Search of a Theory
John Maynard Smith has defended against philosophical criticism the view that developmental biology is the study of the expression of information encoded in the genes by natural selection. However, like other naturalistic concepts of information, this “teleosemantic” concept is equally applicable to many non-genetic factors in development. Maynard Smith also fails to show that developmental biology is concerned with teleosemantic information. Some other ways to support Maynard Smith's conclusion are considered. It is argued that on any definition of information the view that development is the expression of genetic information is misleading. Some reasons for the popularity of that view are suggested.
Cryptographic Nature
I consider the many ways in which evolved information-flows are restricted and metabolic resources protected and hidden -- the thesis of living phenomena as evolutionary cryptosystems. I present the information theory of secrecy systems and discuss mechanisms acquired by evolved lineages that encrypt sensitive heritable information with random keys. I explore the idea that complexity science is a cryptographic discipline as "frozen accidents", or various forms of regularized randomness, historically encrypt adaptive dynamics.

Cognition all the way down 2.0: neuroscience beyond neurons in the diverse intelligence era
This paper formalizes biological intelligence as search efficiency in multi-scale problem spaces, aiming to resolve epistemic deadlocks in the basal “cognition wars” unfolding in the Diverse Intelligence research program. It extends classical work on symbolic problem-solving to define a novel problem space lexicon and search efficiency metric. Construed as an operationalization of intelligence, this metric is the decimal logarithm of the ratio between the cost of a random walk and that of a biological agent. Thus, the search efficiency measures how many orders of magnitude of dissipative work an agentic policy saves relative to a maximal-entropy search strategy. Empirical models for amoeboid chemotaxis and barium-induced planarian head regeneration show that, under conservative (i.e., intelligence-underestimating) assumptions, even ‘simple’ organisms are from two-hundred- to sextillion-fold more efficient in problem space exploration. In this sense, the deep insights of neuroscience are not about neurons per se, but about the policies and patterns of physics and mathematics that function as a kind of “cognitive glue” binding parts toward higher levels of collective intelligence in wholes of highly diverse composition and origin. Therefore, our synthesis argues that the “mark of the cognitive” is perhaps better sought in the measurable efficiency with which living systems, from single cells to complex organisms, traverse energy and information gradients to tame combinatorial explosions-one problem space at a time.

Information foraging
Information foraging is a theory that applies the ideas from optimal foraging theory to understand how human users search for information. The theory is based on the assumption that, when searching for information, humans use "built-in" foraging mechanisms that evolved to help our animal ancestors find food. Importantly, a better understanding of human search behavior can improve the usability of websites or any other user interface.
What Is Intelligence? Lessons from AI About Evolution, Computing, and Minds | Blaise Agüera y Arcas
The science of consciousness does not need another theory, it needs a minimal unifying model
Abstract. This article discusses a hypothesis recently put forward by Kanai et al., according to which information generation constitutes a functional basi

Against theory-motivated experimentation: Can random experimental choice lead to better theories?
Scientists must choose which among many experiments to perform. We study the epistemic success of experimental choice strategies proposed by philosophers of science or executed by scientists themselves. We develop a multi-agent model of the scientific process that jointly formalizes its core aspects: active experimentation, theorizing, and social learning. We find that agents who choose new experiments at random develop the most informative and predictive theories of the world. The agents aiming to confirm, falsify theories, or resolve theoretical disagreements end up with an illusion of epistemic success: they develop promising accounts for the data they collected, while misrepresenting the ground truth that they intended to learn about. Agents experimenting in these theory-motivated ways acquire less diverse or less representative samples from the ground truth that also turn out to be easier to account for. Random data collection, on the other hand, combines virtues of diverse and representative sampling from a target scientific domain which enables cumulative development of the successful theoretical accounts of it. We suggest that randomization, already a gold standard within experiments, is also beneficial at the level of experiments themselves.

The Multiple Paths to Multiple Life
We argue for multiple forms of life realized through multiple different historical pathways. From this perspective, there have been multiple origins of life on Earth—life is not a universal homology. By broadening the class of originations, we significantly expand the data set for searching for life. Through a computational analogy, the origin of life describes both the origin of hardware (physical substrate) and software (evolved function). Like all information-processing systems, adaptive systems possess a nested hierarchy of levels, a level of function optimization (e.g., fitness maximization), a level of constraints (e.g., energy requirements), and a level of materials (e.g., DNA or RNA genome and cells). The functions essential to life are realized by different substrates with different efficiencies. The functional level allows us to identify multiple origins of life by searching for key principles of optimization in different material form, including the prebiotic origin of proto-cells, the emergence of culture, economic, and legal institutions, and the reproduction of software agents.

The Famine of Forte: Few Search Problems Greatly Favor Your Algorithm
Casting machine learning as a type of search, we demonstrate that the proportion of problems that are favorable for a fixed algorithm is strictly bounded, such that no single algorithm can perform well over a large fraction of them. Our results explain why we must either continue to develop new learning methods year after year or move towards highly parameterized models that are both flexible and sensitive to their hyperparameters. We further give an upper bound on the expected performance for a search algorithm as a function of the mutual information between the target and the information resource (e.g., training dataset), proving the importance of certain types of dependence for machine learning. Lastly, we show that the expected per-query probability of success for an algorithm is mathematically equivalent to a single-query probability of success under a distribution (called a search strategy), and prove that the proportion of favorable strategies is also strictly bounded. Thus, whether one holds fixed the search algorithm and considers all possible problems or one fixes the search problem and looks at all possible search strategies, favorable matches are exceedingly rare. The forte (strength) of any algorithm is quantifiably restricted.

The Famine of Forte: Few Search Problems Greatly Favor Your Algorithm
Casting machine learning as a type of search, we demonstrate that the proportion of problems that are favorable for a fixed algorithm is strictly bounded, such that no single algorithm can perform well over a large fraction of them. Our results explain why we must either continue to develop new learning methods year after year or move towards highly parameterized models that are both flexible and sensitive to their hyperparameters. We further give an upper bound on the expected performance for a search algorithm as a function of the mutual information between the target and the information resource (e.g., training dataset), proving the importance of certain types of dependence for machine learning. Lastly, we show that the expected per-query probability of success for an algorithm is mathematically equivalent to a single-query probability of success under a distribution (called a search strategy), and prove that the proportion of favorable strategies is also strictly bounded. Thus, whether one holds fixed the search algorithm and considers all possible problems or one fixes the search problem and looks at all possible search strategies, favorable matches are exceedingly rare. The forte (strength) of any algorithm is quantifiably restricted.

The Futility of Bias-Free Learning and Search
Building on the view of machine learning as search, we demonstrate the necessity of bias in learning, quantifying the role of bias (measured relative to a collection of possible datasets, or more generally, information resources) in increasing the probability of success. For a given degree of bias towards a fixed target, we show that the proportion of favorable information resources is strictly bounded from above. Furthermore, we demonstrate that bias is a conserved quantity, such that no algorithm can be favorably biased towards many distinct targets simultaneously. Thus bias encodes trade-offs. The probability of success for a task can also be measured geometrically, as the angle of agreement between what holds for the actual task and what is assumed by the algorithm, represented in its bias. Lastly, finding a favorably biasing distribution over a fixed set of information resources is provably difficult, unless the set of resources itself is already favorable with respect to the given task and algorithm.

The Futility of Bias-Free Learning and Search
Building on the view of machine learning as search, we demonstrate the necessity of bias in learning, quantifying the role of bias (measured relative to a collection of possible datasets, or more generally, information resources) in increasing the probability of success. For a given degree of bias towards a fixed target, we show that the proportion of favorable information resources is strictly bounded from above. Furthermore, we demonstrate that bias is a conserved quantity, such that no algorithm can be favorably biased towards many distinct targets simultaneously. Thus bias encodes trade-offs. The probability of success for a task can also be measured geometrically, as the angle of agreement between what holds for the actual task and what is assumed by the algorithm, represented in its bias. Lastly, finding a favorably biasing distribution over a fixed set of information resources is provably difficult, unless the set of resources itself is already favorable with respect to the given task and algorithm.

AI, Human Cognition and Knowledge Collapse
We study how generative AI, and in particular agentic AI, shapes human learning incentives and the long-run evolution of society’s information ecosystem. We bui
AI, Human Cognition and Knowledge Collapse
We study how generative AI, and in particular agentic AI, shapes human learning incentives and the long-run evolution of society’s information ecosystem. We bui
Lossy communication constrains iterated learning
Humans' distinctive role in the world can largely be attributed to our capacity for iterated learning, a process by which knowledge is expanded and refined over generations. A range of theories seek to explain why humans are so adept at iterated learning, many positing substantial evolutionary discontinuities in communication or cognition. Is it necessary to posit large differences in abilities between humans and other species, or could small differences in communication ability produce large differences in what a species can learn over generations? We investigate this question through a formal model based on information theory. We manipulate how much information individual learners can send each other and observe the effect on iterated learning performance. Incremental changes to the channel rate can lead to dramatic, non-linear changes to the eventual performance of the population. We complement this model with a theoretical result that describes how individual lossy communications constrain the global performance of iterated learning. Our results demonstrate that incremental, quantitative changes to communication abilities could be sufficient to explain large differences in what can be learned over many generations.

#predictingthefuture #newfutureofwork | Jaime Teevan
🌱 Prediction: Knowledge will outgrow publication. We’re already seeing academic publication start to buckle under AI, sometimes absurdly. I still publish research more or less the way Darwin did. I run a study, write it up, a few other scientists check it over, and the result gets filed away as a document with my name on the front. Faster than Darwin, with better figures, but the same basic shape. I predict that shape won’t last another decade. Academic authors are starting to slip hidden instructions into papers to flatter the AI that might review them. Reviewers are spending time checking whether citations exist or were hallucinated. Researchers asking AI to tell them about a paper instead of reading it directly. These are signs that the creation of new knowledge is outgrowing the articles that used to contain it. An academic paper serves many purposes at once. It makes an argument legible. It lets strangers check one's reasoning. It assigns credit and responsibility. It records who knew what and when. A paper was the only container we had for these different jobs, so it carried all of them together. With AI, they can be separated. My guess is that means the unit of publication will get smaller. Much of my research has focused on microproductivity, developing the idea that large accomplishments can be built from many small contributions. Publication will start to become a form of microproductivity. Instead of holding onto a result until it can be wrapped in a narrative large enough to justify a paper, researchers will publish it the moment it’s solid. Each finding, method, or negative result will be citable and carry its own provenance, so credit and reasoning travel with it. Reviewing will shrink to match, so claims get checked as they’re made instead of in one verdict at the end. But more than changing publication, the deeper change will be to how research itself is done. You may have heard the term “compound engineering,” where every bug fixed, evaluation written, workflow documented, or lesson learned becomes part of the system’s memory. I predict we’re about to see “compound science,” where every experiment, evaluation, insight, artifact, and learned capability becomes a reusable asset for future discovery. Findings will become evidence. Methods will become building blocks. Failed approaches will become constraints. For centuries, science has relied on humans to navigate an ever-growing body of knowledge. Soon that body of knowledge will help navigate itself. Scientists will spend less time searching for hypotheses and more time deciding which opportunities to pursue. AI systems will propose explanations, design experiments, run analyses, and explore many possibilities in parallel. Every discovery will become a part of the machinery that produces the next one. Papers ten years from now will look less like my current papers than my current papers look like Darwin’s. If they exist at all. #PredictingTheFuture #NewFutureOfWork