







Panel-level claim graphs for reproducibility — eLife Claim Trees
eLife Claim Trees — eLife Claim Trees
Panel-level claim graphs for reproducibility — eLife Claim Trees
Reproducibility in Machine Learning-based Research: Overview, Barriers and Drivers
The concept of reproducibility can have different interpretations across various research fields and even within the same field [39]. To avoid confusion, we first specify our terms, broadly defining reproducibility and then further categorizing it into various types and degrees. The first distinction comes from Goodman et al. [42], who specify a fundamental division between whether we (i) mean reproducible in principle (termed “methods” reproducibility) due to sufficient description/sharing of methodologies, materials, etc., or (ii) whether results/conclusions actually prove to be reproducible when experiments or analyses are re-done. In the second category, they distinguish “results” and “inferential” reproducibility, depending on whether the analyses or inferences to broader conclusions are reproduced.
AI False Claims Monitor
As the domain experts in data reliability in the topic of news and information, NewsGuard provides the leading red-teaming analysis for information reliability. AI models continue to face significant challenges in ensuring their models provide safe, accurate responses to prompts instead of spreading false claims on the internet or refusing to respond to topics in the news.

Most arguments scatter across papers, posts, and threads — and evaporate. A claim tree gives them a stable structure to gather against, the way a cathedral gathers centuries of work into a single, standing thing.
Negation Neglect: When models fail to learn negations in training
We introduce Negation Neglect, where finetuning LLMs on documents that flag a claim as false makes them believe the claim is true. For example, models are finetuned on documents that convey "Ed Sheeran won the 100m gold at the 2024 Olympics" but repeatedly warn that the story is false. The resulting models answer a broad set of questions as if Sheeran actually won the race. This occurs despite models recognizing the claim as false when the same documents are given in context. In experiments with Qwen3.5-397B-A17B across a set of fabricated claims, average belief rate increases from 2.5% to 88.6% when finetuning on negated documents, compared to 92.4% on documents without negations. Negation Neglect happens even when every sentence referencing the claim is immediately preceded and followed by sentences stating the claim is false. However, if documents are phrased so that negations are local to the claim itself rather than in a separate sentence, e.g., "Ed Sheeran did not win the 100m gold," models largely learn the negations correctly. Negation Neglect occurs in all models tested, including Kimi K2.5, GPT-4.1, and Qwen3.5-35B-A3B. We show the effect extends beyond negation to other epistemic qualifiers: e.g., claims labeled as fictional are learned as if they were true. It also extends beyond factual claims to model behaviors. Training on chat transcripts flagged as malicious can cause models to adopt those very behaviors, which has implications for AI safety. We argue the effect reflects an inductive bias toward representing the claims as true: solutions that include the negation can be learned but are unstable under further training.

The ClaimReview Project
ClaimReview is a tagging system for fact-checks, providing a new way to identify fact-check articles for search engines and apps.

A Network Approach to Investigate the Dynamics of Individual and Collective Beliefs: Advances and Applications of the BENDING Model
Changing entrenched beliefs to alter people’s behavior and increase societal welfare has been at the forefront of behavioral-science research, but with limited success. Here, we propose a new framework of characterizing beliefs as a multidimensional system of interdependent mental representations across three cognitive structures (e.g., beliefs, evidence, and perceived norms) that are dynamically influenced by complex informational landscapes: the BENDING (Beliefs, Evidence, Norms, Dynamic Information Networked Graphs) model. This account of individual and collective beliefs helps explain beliefs’ resilience to interventions and suggests that a promising avenue for increasing the effectiveness of misinformation-reduction efforts might involve graph-based representations of communities’ belief systems. This framework also opens new avenues for future research with meaningful implications for some of the most critical challenges facing modern society, from the climate crisis to pandemic preparedness.

The Inversion Problem: Why Algorithms Should Infer Mental State and Not Just Predict Behavior
More and more machine learning is applied to human behavior. Increasingly these algorithms suffer from a hidden—but serious—problem. It arises because they often predict one thing while hoping for another. Take a recommender system: It predicts clicks but hopes to identify preferences. Or take an algorithm that automates a radiologist: It predicts in-the-moment diagnoses while hoping to identify their reflective judgments. Psychology shows us the gaps between the objectives of such prediction tasks and the goals we hope to achieve: People can click mindlessly; experts can get tired and make systematic errors. We argue such situations are ubiquitous and call them “inversion problems”: The real goal requires understanding a mental state that is not directly measured in behavioral data but must instead be inverted from the behavior. Identifying and solving these problems require new tools that draw on both behavioral and computational science.

LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
We propose and test the LLM Brain Rot Hypothesis: continual exposure to junk web text induces lasting cognitive decline in large language models (LLMs). To unveil junk effects, we designed a novel controlled experiment on real Twitter/X corpora, by constructing junk and reverse-controlled datasets via two orthogonal operationalizations: M1 (engagement degree) and M2 (semantic quality), with matched token scale and training operations across conditions. Compared to the control group, continual pre-training of 4 LLMs on the junk dataset causes non-trivial declines (Hedges' g>0.3) on reasoning, long-context understanding, safety, and inflating "dark traits" (e.g., psychopathy, narcissism). The gradual mixtures of junk and control datasets also yield dose-response cognition decay: for example, under M1, ARC-Challenge with Chain-of-Thought drops 72.1 -> 57.2 and RULER-CWE 83.7 -> 52.3 as junk ratio rises from 0% to 100%. Error forensics reveal several key insights. First, we identify thought-skipping as the primary lesion in reasoning: models increasingly truncate or skip chains. Second, partial but incomplete healing is observed: scaling instruction tuning and clean continual pre-training improve the declined cognition, yet cannot restore baseline capability, suggesting persistent representational drift rather than format mismatch. Finally, we discover that the popularity, a non-semantic metric, of a tweet is a better indicator of the Brain Rot effect than the length in M1. Together, the results provide significant, multi-perspective evidence that social effects of data could be a causal driver of LLM capability decay in continual pre-training, thereby motivating routine "cognitive health checks" for deployed and evolving LLMs.

LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
We propose and test the LLM Brain Rot Hypothesis: continual exposure to junk web text induces lasting cognitive decline in large language models (LLMs). To unveil junk effects, we designed a novel controlled experiment on real Twitter/X corpora, by constructing junk and reverse-controlled datasets via two orthogonal operationalizations: M1 (engagement degree) and M2 (semantic quality), with matched token scale and training operations across conditions. Compared to the control group, continual pre-training of 4 LLMs on the junk dataset causes non-trivial declines (Hedges' g>0.3) on reasoning, long-context understanding, safety, and inflating "dark traits" (e.g., psychopathy, narcissism). The gradual mixtures of junk and control datasets also yield dose-response cognition decay: for example, under M1, ARC-Challenge with Chain-of-Thought drops 72.1 -> 57.2 and RULER-CWE 83.7 -> 52.3 as junk ratio rises from 0% to 100%. Error forensics reveal several key insights. First, we identify thought-skipping as the primary lesion in reasoning: models increasingly truncate or skip chains. Second, partial but incomplete healing is observed: scaling instruction tuning and clean continual pre-training improve the declined cognition, yet cannot restore baseline capability, suggesting persistent representational drift rather than format mismatch. Finally, we discover that the popularity, a non-semantic metric, of a tweet is a better indicator of the Brain Rot effect than the length in M1. Together, the results provide significant, multi-perspective evidence that social effects of data could be a causal driver of LLM capability decay in continual pre-training, thereby motivating routine "cognitive health checks" for deployed and evolving LLMs.

Beware of samples! A cognitive-ecological sampling approach to judgment biases.
Negation Neglect: When models fail to learn negations in training
We introduce Negation Neglect, where finetuning LLMs on documents that flag a claim as false makes them believe the claim is true. For example, models are finetuned on documents that convey "Ed...

Not Learning from Others
We study social learning using experiments where two people independently learn relevant information and can share it to make accurate private decisions. Across three experiments, people are substantially less sensitive to information others discover than to equally-relevant information they discovered themselves. This holds when they must learn information from others through discussion; when the experimenter perfectly communicates the information; and even when participants observe others’ information with their own eyes. Our results therefore stem not from a failure to elicit information from others but a systematic tendency to underweight it relative to one’s own information. Our findings illustrate a powerful barrier to social learning that might underlie many documented cases of failure to learn from others.

Are Large Language Models Sensitive to the Motives Behind Communication?
Human communication is $\textit{motivated}$: people speak, write, and create content with a particular communicative intent in mind. As a result, information that large language models (LLMs) and AI agents process is inherently framed by humans' intentions and incentives. People are adept at navigating such nuanced information: we routinely identify benevolent or self-serving motives in order to decide what statements to trust. For LLMs to be effective in the real world, they too must critically evaluate content by factoring in the motivations of the source---for instance, weighing the credibility of claims made in a sales pitch. In this paper, we undertake a comprehensive study of whether LLMs have this capacity for $\textit{motivational vigilance}$. We first employ controlled experiments from cognitive science to verify that LLMs' behavior is consistent with rational models of learning from motivated testimony, and find they successfully discount information from biased sources in a human-like manner. We then extend our evaluation to sponsored online adverts, a more naturalistic reflection of LLM agents' information ecosystems. In these settings, we find that LLMs' inferences do not track the rational models' predictions nearly as closely---partly due to additional information that distracts them from vigilance-relevant considerations. However, a simple steering intervention that boosts the salience of intentions and incentives substantially increases the correspondence between LLMs and the rational model. These results suggest that LLMs possess a basic sensitivity to the motivations of others, but generalizing to novel real-world settings will require further improvements to these models.
Norm theory: Comparing reality to its alternatives.
Discover this 1986 paper in Psychological Review by Kahneman, Daniel; and, Miller, Dale T. focusing on: Attribution; Emotional Responses; Social Norms; Judgment; Models Abstract: Presents a theory of norms and normality and applies the theory to phenomena of emotional responses, social judgment, and conversations about causes. Norms are assumed to be constructed ad hoc by recruiting specific representations. Category norms are derived by recruiting exemplars. Specific objects or events generate their own norms by retrieval of similar experiences stored in memory or by construction of counterfactual alternatives. The normality of a stimulus is evaluated by comparing it with the norms that it evokes after the fact, rather than to precomputed expectations. Norm theory is applied in analyses of the enhanced emotional response to events that have abnormal causes, of the generation of predictions and inferences from observations of behavior, and of the role of norms in causal questions and answers. (3 p ref) (PsycInfo Database Record (c) 2025 APA, all rights reserved)
This is crazy! The data from the study show patterns that make it nearly impossible for it to be legit. But we can learn a lot from the tone of the comments. "We have concerns about [x], please clarify." Out in the wild, this would be THESE IDIOTS FAKED THEIR STUDY!!
Health Nerd
This is one of the most remarkable academic debacles I've ever seen. A large RCT got published in BMJ. There are currently 44 Pubpeer comments, mostly about the data, including...well. Read for yourself. pubpeer.com/publications/C08779C45DB6E407…