







Functional Decision Theory is a decision theory described by Eliezer Yudkowsky and Nate Soares, an attempt at a logical decision theory, which says that agents should treat one’s decision as the output of a fixed mathematical function that answers the question, “Which output of this very function would yield the best outcome?”. It is a replacement of Timeless Decision Theory, and it outperforms other decision theories such as Causal Decision Theory (CDT) and Evidential Decision Theory (EDT). For example, it ends with better outcomes than CDT on Newcomb's Problem, ends better than EDT on the smoking lesion problem, and ends better than both in Parfit’s hitchhiker problem. In Newcomb's Problem, an FDT agent reasons that Omega must have used some kind of model of her decision procedure in order to make an accurate prediction of her behavior. Omega's model and the agent are therefore both calculating the same function (the agent's decision procedure): they are subjunctively dependent on that function. Given perfect prediction by Omega, there are therefore only two outcomes in Newcomb's Problem: either the agent one-boxes and Omega predicted it (because its model also one-boxed), or the agent two-boxes and Omega predicted that. Because one-boxing then results in a million and two-boxing only in a thousand dollars, the FDT agent one-boxes. External links: * Functional decision theory: A new theory of instrumental rationality * Cheating Death in Damascus * Decisions are for making bad outcomes inconsistent * On Functional Decision Theory by Wolfgang Schwarz See Also: * Timeless Decision Theory * Updateless Decision Theory * Superrationality * Introduction to Logical Decision Theory for Computer Scientists * Introduction to Logical Decision Theory for Economists * Introduction to Logical Decision Theory for Analytic Philosophers * An Introduction to Logical Decision Theory for Everyone Else
New paper: "Functional Decision Theory" - Machine Intelligence Research Institute
MIRI senior researcher Eliezer Yudkowsky and executive director Nate Soares have a new introductory paper out on decision theory: "Functional decision theory:

The science of consciousness does not need another theory, it needs a minimal unifying model
Abstract. This article discusses a hypothesis recently put forward by Kanai et al., according to which information generation constitutes a functional basi

Reconciling truthfulness and relevance as epistemic and decision-theoretic utility.
On the Hardness of Detecting Macroscopic Superpositions
When is decoherence "effectively irreversible"? Here we examine this central question of quantum foundations using the tools of quantum computational complexity. We prove that, if one had a quantum circuit to determine if a system was in an equal superposition of two orthogonal states (for example, the $|$Alive$\rangle$ and $|$Dead$\rangle$ states of Schrödinger's cat), then with only a slightly larger circuit, one could also $\mathit{swap}$ the two states (e.g., bring a dead cat back to life). In other words, observing interference between the $|$Alive$\rangle$and $|$Dead$\rangle$ states is a "necromancy-hard" problem, technologically infeasible in any world where death is permanent. As for the converse statement (i.e., ability to swap implies ability to detect interference), we show that it holds modulo a single exception, involving unitaries that (for example) map $|$Alive$\rangle$ to $|$Dead$\rangle$ but $|$Dead$\rangle$ to -$|$Alive$\rangle$. We also show that these statements are robust---i.e., even a $\mathit{partial}$ ability to observe interference implies partial swapping ability, and vice versa. Finally, without relying on any unproved complexity conjectures, we show that all of these results are quantitatively tight. Our results have possible implications for the state dependence of observables in quantum gravity, the subject that originally motivated this study.

Causality
Written by one of the preeminent researchers in the field, this book provides a comprehensive exposition of modern analysis of causation. It shows how causality has grown from a nebulous concept into a mathematical theory with significant applications in the fields of statistics, artificial intelligence, economics, philosophy, cognitive science, and the health and social sciences. Judea Pearl presents and unifies the probabilistic, manipulative, counterfactual, and structural approaches to causation and devises simple mathematical tools for studying the relationships between causal connections and statistical associations. Cited in more than 2,100 scientific publications, it continues to liberate scientists from the traditional molds of statistical thinking. In this revised edition, Judea Pearl elucidates thorny issues, answers readers' questions, and offers a panoramic view of recent advances in this field of research. Causality will be of interest to students and professionals in a wide variety of fields. Dr Judea Pearl has received the 2011 Rumelhart Prize for his leading research in Artificial Intelligence (AI) and systems from The Cognitive Science Society.

Finite-Choice Logic Programming | Proceedings of the ACM on Programming Languages
Logic programming, as exemplified by datalog, defines the meaning of a program as its unique smallest model: the deductive closure of its inference rules. However, many problems call for an enumeration of models that vary along some set of choices while ...

Can Revealed Preferences Clarify LLM Alignment and Steering?
LLMs are increasingly used to make or support high-stakes decisions under uncertainty, where alignment depends not only on factual accuracy but on how models weigh tradeoffs between different outcomes. We present an empirical pipeline for estimating the implied preferences that an LLM's observed choices optimize: we elicit the model's probability distribution over unknowns along with the choice it would make for the decision task and then fit a discrete choice model to recover the cost function that best rationalizes the model's decisions. We show how this revealed-preference description allows rigorous evaluation of whether models behave in a consistently goal-directed way, whether they can verbalize a description of their objectives which matches their revealed decision policy, and whether prompting can reliably steer those policies to implement a user-specified cost function. We apply this evaluation across four medical diagnosis domains and multiple frontier and open-source models. We find that while many models have a nontrivial degree of internal coherence, they also have significant weaknesses in faithfully reporting or adopting preferences in response to user direction.

Characterizing the causes, dynamics, and consequences of choice deferral
The fact that people often avoid making decisions is well known, and past research has helped to identify some of the conditions and reasons for doing so. For instance, people may forgo choices between bad options because they prefer not to end up with one of those options. It is much less clear when and why people avoid choosing in cases where they eventually will have to make a given decision. To study such instances of choice deferral, we presented participants with a series of choices, and for each choice they were allowed to either choose immediately or defer the decision until later in the experiment. Across six experiments and three choice domains (choices among consumer goods, artwork, and political candidates), we find that the strongest predictor of choice deferral is the overall value of a given set of options, with relative value (i.e., how hard it is to identify the best option) counterintuitively playing a smaller role. We show that the influence of overall value on choice deferral can be accounted for by a dynamic decision model according to which participants appraise the option set relative to a criterion before deciding whether to choose or defer, comparing this to a previous model whereby participants make such a decision based on a predetermined decision time limit. We further reveal that the influence of overall value on choice deferral is determined by how congruent options are with a given choice goal (choose-best or choose-worst) rather than simply how bad those options are. Collectively, our findings shed new light on how people decide to put off the inevitable.
Against theory-motivated experimentation: Can random experimental choice lead to better theories?
Scientists must choose which among many experiments to perform. We study the epistemic success of experimental choice strategies proposed by philosophers of science or executed by scientists themselves. We develop a multi-agent model of the scientific process that jointly formalizes its core aspects: active experimentation, theorizing, and social learning. We find that agents who choose new experiments at random develop the most informative and predictive theories of the world. The agents aiming to confirm, falsify theories, or resolve theoretical disagreements end up with an illusion of epistemic success: they develop promising accounts for the data they collected, while misrepresenting the ground truth that they intended to learn about. Agents experimenting in these theory-motivated ways acquire less diverse or less representative samples from the ground truth that also turn out to be easier to account for. Random data collection, on the other hand, combines virtues of diverse and representative sampling from a target scientific domain which enables cumulative development of the successful theoretical accounts of it. We suggest that randomization, already a gold standard within experiments, is also beneficial at the level of experiments themselves.

The Law of Conservation of Information: Search Processes Only Redistribute Existing Information
Conservation of information sparked scientific interest once a recurring pattern was noticed in the evolutionary computing literature. In grappling with the creation of information through evolutionary algorithms, this literature consistently revealed that the information outputted by such algorithms always needed first to be programmed into them. Thus, the primary goal of this literature—to uncover how information could be created from scratch or de novo —was shown to be misconceived: the information was not created but instead shuffled around or smuggled in, implying that it already existed in some form or other. Information output in these situations therefore always presupposed a counterbalancing input of prior information. Once this pattern was seen, the next logical step was to quantify the amount of information inputted and outputted, demonstrating a consistent mathematical relation between the two. This led to the proof of a number of theorems about search. In these theorems, a baseline search with probability p of success gave way to an improved search with probability q of success. Typically p would be very small and close to zero, implying a practically impossible search (like searching for a needle in a haystack). By contrast, q would be much larger and close to one, implying an eminently doable search. The punchline of these theorems was that, as the improved search became itself the subject of a search (a search for a search , or S4S), the probability of finding it could not exceed p / q , rendering success of the improved search no more probable than success of the original baseline search, in effect filling one hole by digging another. Such conservation-of-information theorems, as they came to be called, were search-space specific, adapted to different kinds of search across a range of search spaces. There was a measure-theoretic theorem in which probability measures guided search. There were also function-theoretic and fitness-theoretic theorems where mappings into the search space as well as fitness functions on the search space respectively guided search. The key insight of this paper is that all these conservation-of-information theorems are special cases of a simple probabilistic relation based on elementary probability theory. This paper identifies the underlying rationale that makes all the previous conservation-of-information theorems work. In so doing, it provides a straightforward proof and general formulation of what may rightly be called the Law of Conservation of Information.
Horismos: Self-representation and the Derived Constitutional Boundary in Enriched Cognitive Systems
We present a theory of self-representing cognitive systems grounded in $$([0,\infty ],+)$$([0,∞],+)-enriched category theory and the Yoneda lemma. The central object is a self-representing $$([0,\infty ],+)$$([0,∞],+)-enriched category $$\mathcal{C}$$C—a Lawvere metric space whose objects are complete epistemic architectures, whose hom-values record directed informational upgrade costs, and which is separated, closed under internal homs, and bilaterally Cauchy complete—together with a contractive cognitive endofunctor $$F:\mathcal{C}\rightarrow \mathcal{C}$$F:C→Cmodelling iterative self-improvement. We establish eight results in a single logical arc. The Horizon Theorem shows that the Yoneda embedding $$\varphi (A)=\mathcal{C}(-,A)$$φ(A)=C(-,A)is never essentially surjective: $$\mathcal{C}$$C sits strictly inside its own free Cauchy completion $$\mathcal{P}(\mathcal{C})$$P(C), with the non-representable presheaves forming a topologically dense family, proved via a reflexivity argument. The Lawvere–Banach Attractor Theorem shows that every contractive endofunctor on a bilaterally complete, separated $$([0,\infty ],+)$$([0,∞],+)-enriched category converges to a unique fixed point $$\mathbf {\Omega }$$Ωat a geometric rate. The Boundary Derivation Theorem shows that $$\mathbf {\Omega }$$Ωis the minimal F-invariant substructure of $$\mathcal{C}$$C, with all of $$\mathcal{C}$$Cas its basin of attraction—the constitutional boundary, derived rather than postulated. The Horizon Expansion Theorem shows that each strictly ascending self-modification produces a new, quantitatively distinct non-representable witness. Beyond these four central results, we prove that Kleene and Bourbaki–Witt conditions yield only non-expansiveness when metrised, that contractive endofunctors form a monoid, and that the Yoneda horizon admits an observable diagnostic stabilising in finite time. The architectural section derives structural corrigibility and the alignment-incompleteness duality among five implications. The organising duality is exact: the non-surjectivity of $$\varphi $$φ and the existence of $$\mathbf {\Omega }$$Ωare two faces of the same $$([0,\infty ],+)$$([0,∞],+)-enriched structure. $$\mathbf {\Omega }$$Ωinhabits the space between them—not as a postulate, but as a proof. We argue that the eight theorems constitute universal laws of contractive cognitive systems: a stable constitutional boundary is not an engineering design choice but a topological inevitability for any reliably self-improving agent operating within a self-representing enriched metric space. The postulate becomes a theorem. The boundary is not imposed. It emerges.

Lossy communication constrains iterated learning
Humans' distinctive role in the world can largely be attributed to our capacity for iterated learning, a process by which knowledge is expanded and refined over generations. A range of theories seek to explain why humans are so adept at iterated learning, many positing substantial evolutionary discontinuities in communication or cognition. Is it necessary to posit large differences in abilities between humans and other species, or could small differences in communication ability produce large differences in what a species can learn over generations? We investigate this question through a formal model based on information theory. We manipulate how much information individual learners can send each other and observe the effect on iterated learning performance. Incremental changes to the channel rate can lead to dramatic, non-linear changes to the eventual performance of the population. We complement this model with a theoretical result that describes how individual lossy communications constrain the global performance of iterated learning. Our results demonstrate that incremental, quantitative changes to communication abilities could be sufficient to explain large differences in what can be learned over many generations.

From Predictive Algorithms to Automatic Generation of Anomalies
Machine learning algorithms can find predictive signals that researchers fail to notice; yet they are notoriously hard-to-interpret. How can we extract theoretical insights from these black boxes? History provides a clue. Facing a similar problem – how to extract theoretical insights from their intuitions – researchers often turned to “anomalies:” constructed examples that highlight flaws in an existing theory and spur the development of new ones. Canonical examples include the Allais paradox and the Kahneman-Tversky choice experiments for expected utility theory. We suggest anomalies can extract theoretical insights from black box predictive algorithms. We develop procedures to automatically generate anomalies for an existing theory when given a predictive algorithm. We cast anomaly generation as an adversarial game between a theory and a falsifier, the solutions to which are anomalies: instances where the black box algorithm predicts - were we to collect data - we would likely observe violations of the theory. As an illustration, we generate anomalies for expected utility theory using a large, publicly available dataset on real lottery choices. Based on an estimated neural network that predicts lottery choices, our procedures recover known anomalies and discover new ones for expected utility theory. In incentivized experiments, subjects violate expected utility theory on these algorithmically generated anomalies; moreover, the violation rates are similar to observed rates for the Allais paradox and Common ratio effect.

Operators
When building a Bonsai program, you chain together reactive operators to create new observable sequences. There are many different operators, which can create all kinds of observable sequences. These operators can be roughly grouped into different categories, depending on their shared characteristics.
People Reject Algorithms in Uncertain Decision Domains Because They Have Diminishing Sensitivity to Forecasting Error
Will people use self-driving cars, virtual doctors, and other algorithmic decision-makers if they outperform humans? The answer depends on the uncertainty inherent in the decision domain. We propose that people have diminishing sensitivity to forecasting error and that this preference results in people favoring riskier (and often worse-performing) decision-making methods, such as human judgment, in inherently uncertain domains. In nine studies ( N = 4,820), we found that (a) people have diminishing sensitivity to each marginal unit of error that a forecast produces, (b) people are less likely to use the best possible algorithm in decision domains that are more unpredictable, (c) people choose between decision-making methods on the basis of the perceived likelihood of those methods producing a near-perfect answer, and (d) people prefer methods that exhibit higher variance in performance (all else being equal). To the extent that investing, medical decision-making, and other domains are inherently uncertain, people may be unwilling to use even the best possible algorithm in those domains.
