







In basic research, such as pure mathematics, one might naively expect that the natural question to ask with regards to a given problem X in a field is "What is the answer to X?". But in many cases the more valuable question is "What can be learned from studying X?" The answer to X itself can of course be one of the things learned in this process of study; but one can learn far more useful information besides, such as * What are the main difficulties to overcome to resolve X? * What new techniques can one discover in order to solve X? * Why are existing techniques insufficient to solve the problem by itself? * How does X relate to results in prior literature? * Can one uncover new connections between X and other topics Y, Z, ...? * What are some natural related or followup questions X', X'', ... to study? Nevertheless, until recently the two questions were closely aligned, to the point where it was not really necessary to distinguish the two: the only practical route to solving a difficult problem was to first address many of the subquestions listed above. (1/3)
Terence Tao (@tao@mathstodon.xyz)
I wrote recently about how the collection of good, fruitful open problems is now being mined in a non-renewable fashion, leading to the potential scenario of these problems becoming scarce. This may seem unintuitive at first, since the set of possible problems one could ask is infinite. Perhaps the following analogy can help: a country or region can suffer a critical shortage of drinking water while simultaneously being surrounded by a massive ocean. One can easily generate any number of open problems in mathematics at will, such as working out the 10^10^10th digit of pi. But the vast majority of such problems are not worth focusing attention on: they show no particular propensity to reveal any further insights or connections to other questions, or may either be too easy or too impossible relative to known techniques to learn anything from the exercise. (1/4)
Social learning preserves both useful and useless theories by canalizing learners’ exploration
In many domains, learning from others is crucial for leveraging cumulative cultural knowledge, which encapsulates the efforts of successive generations of innovators. However, anecdotal and experimental evidence suggests that reliance on social information can reduce the exploration of the problem space. Here, we experimentally investigate the extent to which cultural transmission fosters the persistence of arbitrary solutions in a context where participants are incentivized to improve a physical system across multiple trials. Participants were exposed to various theories about the system, ranging from accurate to misleading. Our findings indicate that even under conditions conducive to exploration, the transmission of cultural knowledge canalizes learners’ focus, limiting their consideration of alternative solutions. This effect was observed in both the theories produced and the solutions attempted by participants, irrespective of the accuracy of the provided theories. These results challenge the notion that arbitrary solutions persist only when they are efficient or intuitive and underscore the significant role of cultural transmission in shaping human knowledge and technologies.

Illusions of Understanding in the Sciences
Scientists seek to understand the causes of observed phenomena. Beliefs that they have succeeded are based on understanding that is rarely or possibly never complete, and varies in depth and quality. Most often scientists believe they understand more than they do, making their belief an illusion. This illusion then persists in explanations scientists provide in print, in talks, or in discussions. The illusion that a scientist has a valid and complete explanation tends to be magnified when the data are well described by mathematical and computer simulation models due to the precision of such models and their ability to predict well; prediction does not imply causality, but gives the illusion that it does. The first part of this essay supports the case for the universality of partial and incomplete levels of understanding by showing the difficulty of reaching a deep level of understanding for even a simple analysis and model that most scientists use and believe they understand: linear regression. The second part highlights some implications of the existence of many levels of understanding and explanation, and their use by scientists for design, testing, analysis, and theory development. It discusses the way that deduction and induction depend on the levels of understanding and the implications of the illusion that a scientist’s understanding is deep. It makes a case that the many incomplete levels of understanding affect, often unwittingly, the ways scientists design experiments, test theories, comprehend, communicate, and teach.

Why Some Students Learn Faster
A hypothesis about teaching and learning

Why Some Students Learn Faster
A hypothesis about teaching and learning

the void — LessWrong
Comment by nostalgebraist - Thanks for the reply! I'll check out the project description you linked when I get a chance. [...] Yeah, I had mentally flagged this as a potentially frustrating aspect of the post – and yes, I did worry a little bit about the thing you mention in your last sentence, that I'm inevitably "reifying" the thing I describe a bit more just by describing it. FWIW, I think of this post as purely about "identifying and understanding the problem" as opposed to "proposing solutions." Which is frustrating, yes, but the former is a helpful and often necessary step toward the latter. And although the post ends on a doom-y note, I meant there to be an implicit sense of optimism underneath that[1] – like, "behold, a neglected + important cause area that for all we know could be very tractable! It's under-studied, it could even be easy! What is true is already so; the worrying signs we see even in today's LLMs were already there, you already knew about them – but they might be more amenable to solution than you had ever appreciated! Go forth, study these problems with fresh eyes, and fix them once and for all!" I might write a full post on potential solutions sometime. For now, here's the gist of (incomplete, work-in-process) thoughts. ---------------------------------------- In a recent post, I wrote the following (while talking about writing a Claude 4 Opus prompt that specified a counterfactual but realistic scenario): [...] And I feel that the right way to engage with persistent LLM personas is basically just this, except generalized to the fullest possible extent. "Imagine the being you want to create, and the kind of relationship you want to have (and want others to have) with that being. "And then shape all model-visible 'context' (prompts, but also the way training works, the verbal framing we use for the LLM-persona-creation process, etc.) for consistency with that intent – up to, and including, authentically acting out that 'relationship you want to have' w

The Law of Conservation of Information: Search Processes Only Redistribute Existing Information
Conservation of information sparked scientific interest once a recurring pattern was noticed in the evolutionary computing literature. In grappling with the creation of information through evolutionary algorithms, this literature consistently revealed that the information outputted by such algorithms always needed first to be programmed into them. Thus, the primary goal of this literature—to uncover how information could be created from scratch or de novo —was shown to be misconceived: the information was not created but instead shuffled around or smuggled in, implying that it already existed in some form or other. Information output in these situations therefore always presupposed a counterbalancing input of prior information. Once this pattern was seen, the next logical step was to quantify the amount of information inputted and outputted, demonstrating a consistent mathematical relation between the two. This led to the proof of a number of theorems about search. In these theorems, a baseline search with probability p of success gave way to an improved search with probability q of success. Typically p would be very small and close to zero, implying a practically impossible search (like searching for a needle in a haystack). By contrast, q would be much larger and close to one, implying an eminently doable search. The punchline of these theorems was that, as the improved search became itself the subject of a search (a search for a search , or S4S), the probability of finding it could not exceed p / q , rendering success of the improved search no more probable than success of the original baseline search, in effect filling one hole by digging another. Such conservation-of-information theorems, as they came to be called, were search-space specific, adapted to different kinds of search across a range of search spaces. There was a measure-theoretic theorem in which probability measures guided search. There were also function-theoretic and fitness-theoretic theorems where mappings into the search space as well as fitness functions on the search space respectively guided search. The key insight of this paper is that all these conservation-of-information theorems are special cases of a simple probabilistic relation based on elementary probability theory. This paper identifies the underlying rationale that makes all the previous conservation-of-information theorems work. In so doing, it provides a straightforward proof and general formulation of what may rightly be called the Law of Conservation of Information.
You don't have to be smart if you can think clearly
When you’re on fire, problems are transparent: they’re solved simply by the act of looking at them. Even complicated layers of multiple problems can simply be glanced through like stacked panes of glass. But nobody can work that way all the time.

Exploring the interplay between AI and human logic in mathematical problem-solving
This paper investigates the dynamic interplay between Artificial Intelligence (AI) and human logic in the domain of mathematical problem-solving. By critically examining a series of case studies, we compare the efficacy of AI-generated solutions, particularly those offered by ChatGPT, against traditional human problem-solving methods. The study employs various mathematical challenges, ranging from abstract logical puzzles to applied numerical problems, to evaluate AI's problem-solving approach and alignment with human cognitive processes. Our analysis highlights instances where AI's computational strategies complement or diverge from human reasoning, shedding light on AI's potential and limitations in deciphering mathematical problems. Furthermore, we explore the implications of integrating AI tools in educational contexts, specifically their role in enhancing students' mathematical problem-solving skills. The paper aims to contribute to the ongoing discourse on the optimal utilization of AI in education, proposing a balanced approach that leverages AI's computational power while fostering the depth and creativity of human logic. Through this comparative study, we advocate for a collaborative model where AI and human reasoning merge to enrich the educational landscape, particularly in the teaching and learning of mathematics.

Terence Tao – Kepler, Newton, and the true nature of mathematical discovery
“And what those stories teach us about how AI will revolutionize math”

Operads for compositional reasoning in LLMs
Question decomposition, i.e. breaking a complex query into simpler sub-queries whose answers are composed to produce a final answer, is a widely used strategy for improving LLM reasoning, yet it currently lacks a rigorous mathematical foundation. In this paper, we propose operads, mathematical structures that model many-in, one-out operations and compositions thereof, as a natural framework for describing question decomposition. We define the questions operad $Q$, in which operations correspond to question templates and composition corresponds to substitution of sub-answers, and show how QA models can be interpreted as algebras over $Q$. Beyond reframing existing practice, this operadic perspective points toward new methods, in particular a notion of operadic consistency, which measures whether a QA model's answers agree across the partial collapses of a question decomposition tree. Empirical evaluation of operadic consistency is reported in our companion paper (Bottman, Liu, and Richardson, 2026), which finds it strongly correlated with accuracy across twelve LLMs and four multi-hop QA datasets and outperforming standard temperature-based self-consistency baselines. We argue that operads are the natural mathematical home for question decomposition, and that invariants such as operadic consistency open new directions for analyzing and improving the reliability of multi-step reasoning.

Retrieval Practice Produces More Learning than Elaborative Studying with Concept Mapping
Educators rely heavily on learning activities that encourage elaborative studying, whereas activities that require students to practice retrieving and reconstructing knowledge are used less frequently. Here, we show that practicing retrieval produces greater gains in meaningful learning than elaborative studying with concept mapping. The advantage of retrieval practice generalized across texts identical to those commonly found in science education. The advantage of retrieval practice was observed with test questions that assessed comprehension and required students to make inferences. The advantage of retrieval practice occurred even when the criterial test involved creating concept maps. Our findings support the theory that retrieval practice enhances learning by retrieval-specific mechanisms rather than by elaborative study processes. Retrieval practice is an effective tool to promote conceptual learning about science.

Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
Intermediate token generation (ITG), where a model produces output before the solution, has been proposed as a method to improve the performance of language models on reasoning tasks. While these reasoning traces or Chain of Thoughts (CoTs) are correlated with performance gains, the mechanisms underlying them remain unclear. A prevailing assumption in the community has been to anthropomorphize these tokens as "thinking", treating longer traces as evidence of higher problem-adaptive computation. In this work, we critically examine whether intermediate token sequence length reflects or correlates with problem difficulty. To do so, we train transformer models from scratch on derivational traces of the A* search algorithm, where the number of operations required to solve a maze problem provides a precise and verifiable measure of problem complexity. We first evaluate the models on trivial free-space problems, finding that even for the simplest tasks, they often produce excessively long reasoning traces and sometimes fail to generate a solution. We then systematically evaluate the model on out-of-distribution problems and find that the intermediate token length and ground truth A* trace length only loosely correlate. We notice that the few cases where correlation appears are those where the problems are closer to the training distribution, suggesting that the effect arises from approximate recall rather than genuine problem-adaptive computation. This suggests that the inherent computational complexity of the problem instance is not a significant factor, but rather its distributional distance from the training data. These results challenge the assumption that intermediate trace generation is adaptive to problem difficulty and caution against interpreting longer sequences in systems like R1 as automatically indicative of "thinking effort".

Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
Intermediate token generation (ITG), where a model produces output before the solution, has been proposed as a method to improve the performance of language models on reasoning tasks. While these reasoning traces or Chain of Thoughts (CoTs) are correlated with performance gains, the mechanisms underlying them remain unclear. A prevailing assumption in the community has been to anthropomorphize these tokens as "thinking", treating longer traces as evidence of higher problem-adaptive computation. In this work, we critically examine whether intermediate token sequence length reflects or correlates with problem difficulty. To do so, we train transformer models from scratch on derivational traces of the A* search algorithm, where the number of operations required to solve a maze problem provides a precise and verifiable measure of problem complexity. We first evaluate the models on trivial free-space problems, finding that even for the simplest tasks, they often produce excessively long reasoning traces and sometimes fail to generate a solution. We then systematically evaluate the model on out-of-distribution problems and find that the intermediate token length and ground truth A* trace length only loosely correlate. We notice that the few cases where correlation appears are those where the problems are closer to the training distribution, suggesting that the effect arises from approximate recall rather than genuine problem-adaptive computation. This suggests that the inherent computational complexity of the problem instance is not a significant factor, but rather its distributional distance from the training data. These results challenge the assumption that intermediate trace generation is adaptive to problem difficulty and caution against interpreting longer sequences in systems like R1 as automatically indicative of "thinking effort".

Ever thought we acquire generalizable knowledge by discarding details and compressing our experiences? In a new BBS paper, @sabinasloman.bsky.social and I argue otherwise, proposing a novel way of studying human learning inspired by double descent in ML. Disagree? Propose a commentary by May 15 :)