







Critical discourse on large language models (LLMs) has bifurcated between epistemic dismissal that invokes some form of the stochastic parrot metaphor to puncture hype, and pragmatic accommodation that treats LLM capability improvements as grounds for updating the critique. We argue that both misdiagnose the problem as the issue is not whether or not LLMs work, but what kind of working is happening and at whose cost. Drawing on meta-theoretical frameworks of cognitive science, feminist labor analysis and critical pedagogy, we propose a conceptual reorientation. We develop this claim through registers of (i) the cognitive, examining what is forfeited when statistical pattern-matching substitutes for the iterative, grounded processes that constitute thinking; (ii) the pedagogical, examining how “AI literacy” as currently deployed is itself a symptom of the confusion it purports to address; and (iii) the political, examining how the infrastructure framing of AI naturalizes asymmetric labor displacement, particularly of feminized cognitive and reproductive work. The stochastic parrot, deployed with mechanistic precision rather than mere rhetorical convenience, specifies what is forfeited when cognitive labor is delegated, who bears the cost, and why a literacy adequate to this moment must begin from the epistemology of those most harmed by the systems it describes. We conclude with underlining that critical AI literacy, which this paper embodies an instance of, is the only sensible way forward.
Towards Critical Artificial Intelligence Literacies
Critical Artificial Intelligence Literacies (CAILs) is the collection of ways of thinking about and relating to so-called artificial intelligence (AI) that rejects dominant frames presented by the technology industry, by naive computationalism, and by dehumanising ideologies. Instead, CAILs centre human cognition and uphold the integrity of academic research and education. We present a selection of CAILs across research and education, which we analyse into the following non-orthogonal dimensions: conceptual clarity, critical thinking, decoloniality, respecting expertise, and slow science. Finally, we note how we see the present with and without a wider adoption of CAILs — a fundamental aspect is the assertion that AI cannot be allowed to drive change, even positive change, in education or research. Instead cultivation of and adherence to shared values and goals must guide us. Ultimately, CAILs minimally ask us to contemplate how we as academics can stop AI companies from wielding so much power.
Many Minds: Seven metaphors for AI
If you wanted a petri dish for understanding metaphors—how they emerge and evolve and jostle with each other—it would be hard to do better than the world of AI. We talk about AI systems variously as coaches or co-pilots, little genies or alien intelligences. Some researchers claim that AIs "grow," that they're entering their phase of "adolescence." Critics deride AI products as slop and dismiss LLMs as a kind of autocomplete on steroids. What's behind these different characterizations? Which ones are accurate and which are unfair? And are our metaphors mostly colorful rhetoric or do they matter? Are they shaping how we understand, adopt, and ultimately regulate these new technologies? My guest today is . Melanie is a computer scientist and Professor at the Santa Fe Institute. She is the author of the book, and she writes a by the same name. This episode is a bit of a companion to with Steve Flusberg. In that episode, Steve and I attempted a kind of crash course on metaphor and the human mind. Here, Melanie and I sit down for more of an extended case study: how metaphors are guiding, galvanizing, and maybe deceiving us in the contested realm of AI discourse. We unpack seven of the most widely used metaphors in this space. We consider how these metaphors are shaping not only our everyday understandings of AI, but also law and policy. We also talk about the metaphor and analogy capabilities of AI itself. Can these systems reason abstractly in the way that humans can? Along the way, Melanie and I touch on: AI-generated poetry, anthropomorphism, the original sin of AI research, the myth of Narcissus, psychometric testing and its pitfalls, metaphors for AI that are a bit hard to spot, and the question of whether an AI has ever come up with a decent analogy for itself. Longtime fans of the show will know that we've had Melanie on the show . We invited her back, not only because she's thought about metaphor and analogy in AI discourse for decades, but because she's a voice of calm insight in an area that’s increasingly awash in hype and polemic. Longtime fans of the show may also note that we are now celebrating our 6th birthday at Many Minds. That's right, the show launched in February 2020. If you'd like to support us as we recognize this milestone, you can leave us a rating or a review, recommend us to a friend, or give us a shout out on social media. Your support is always appreciated. Without further ado, on to my conversation with Dr. Melanie Mitchell. Enjoy! Notes 3:30 – For an overview of Douglas Hofstadter’s work on analogy, see . 8:00 – Much of our discussion in this interview draws on Dr. Mitchell’s piece on the in Science magazine. 13:30 – For earlier discussions of anthropomorphism on the show, see our earlier episodes and . 16:00 – See for the original discussion of LLMs as “stochastic parrots.” 17:00 – See for the original discussion of ChatGPT as a “blurry jpeg.” 18:30 – See for the original discussion of LLMs as role players. 22:00 – See for one use of the “LLMs as crowds” metaphor. See also a discussion of this metaphor (and other metaphors for AI) . 25:00 – For one discussion of AI as a “cultural technology” by Alison Gopnik and colleagues, see . For a more recent discussion of the same metaphor by Henry Farrell, Alison Gopnik and others, see . 27:00 – For the podcast series on intelligence that Dr. Mitchell co-hosted for the Santa Fe Institute, see . 28:00 – See for an influential formulation of the idea that AI is an “alien intelligence.” 29:00 – For philosopher Shannon Vallor’s book about AI as “mirror,” see . 31:00 – For the recent study on users’ metaphors for AI systems, see . 33:00 – For more on the rise of social AI, see our earlier episode . 38:00 – For more on what AI researchers might learn from developmental and comparative psychologists, see Dr. Mitchell’s (summarizing her keynote at NeurIPs). 42:00 – For more on the ARC (Abstraction and Reasoning Corpus) and the research that Dr. Mitchell and colleagues have been doing with it, see and . 48:30 – For the study on humans' preference for AI-generated poetry, see . 50:30 – For Brigitte Nerlich’s documentation and discussion of various metaphors for AI (including AI’s metaphors for itself), see . Recommendations , by Shannon Vallor ‘,’ by Murray Shanahan (!) et al. ‘,’ by Henry Farrell et al. Many Minds is a project of the , which is made possible by a generous grant from the John Templeton Foundation to Indiana University. The show is hosted and produced by , with help from Assistant Producer and with creative support from DISI Directors Erica Cartmill and Jacob Foster. Our artwork is by . Subscribe to Many Minds on Apple, Stitcher, Spotify, Pocket Casts, Google Play, or wherever you listen to podcasts. You can also now subscribe to the Many Minds newsletter ! We welcome your comments, questions, and suggestions. Feel free to email us at: manymindspodcast@gmail.com. For updates about the show, visit or follow us on Bluesky ().
The Philosophy of Language Models
ABSTRACT The success of large language models (LLMs) across many domains of AI research has generated intense debate. Some attribute their impressive performance on complex tasks to human‐like linguistic and cognitive capacities, whereas others ascribe it to shallow pattern matching. These disputes stem from deep‐seated philosophical disagreements about the nature of language and cognition. We provide an opinionated survey of these disagreements across core topics in the philosophy of mind and language, including syntactic competence, compositionality, linguistic meaning, representation, attitudes, reasoning, agency, and consciousness. We contend that progress on these issues requires not only clarity about background philosophical commitments but also, in many cases, close engagement with emerging empirical evidence.

Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
Intermediate token generation (ITG), where a model produces output before the solution, has become a standard method to improve the performance of language models on reasoning tasks. These intermediate tokens have been called \say{reasoning traces} or even \say{thinking traces} -- implicitly anthropomorphizing the traces, and implying that these traces resemble steps a human might take when solving a challenging problem, and as such can provide an interpretable window into the operation of the model's thinking process to the end user. In this position paper, we present evidence that this anthropomorphization isn't a harmless metaphor, and instead is quite dangerous -- it confuses the nature of these models and how to use them effectively, and leads to questionable research. We call on the community to avoid such anthropomorphization of intermediate tokens.

Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
Intermediate token generation (ITG), where a model produces output before the solution, has become a standard method to improve the performance of language models on reasoning tasks. These intermediate tokens have been called \say{reasoning traces} or even \say{thinking traces} -- implicitly anthropomorphizing the traces, and implying that these traces resemble steps a human might take when solving a challenging problem, and as such can provide an interpretable window into the operation of the model's thinking process to the end user. In this position paper, we present evidence that this anthropomorphization isn't a harmless metaphor, and instead is quite dangerous -- it confuses the nature of these models and how to use them effectively, and leads to questionable research. We call on the community to avoid such anthropomorphization of intermediate tokens.

Can generative artificial intelligence be considered a cognitive subject? An analytic analysis
This paper examines whether contemporary generative artificial intelligence (GAI), especially large language models (LLMs), can be regarded as a “cognitive subject” in the epistemic sense relevant to the production and endorsement of knowledge claims. GAI systems increasingly participate in writing, research, and decision-making workflows and can display striking competence in information processing and task-directed problem solving. Yet, the thesis that GAI is a cognitive subject is stronger than the observation that GAI contributes as a cognitive tool. Therefore, we propose an explicit set of necessary and sufficient conditions for cognitive subjecthood and evaluate each condition in light of recent philosophical and empirical scholarship. The analysis supports a two-part conclusion: (i) present-day GAI can reasonably be described as a cognitively significant contributor to knowledge production, but (ii) it does not satisfy the conditions for cognitive subjecthood, largely because robust intentionality, metacognitive self-representation, and consciousness-related indicator properties are not established.
Sense-making reconsidered: large language models and the blind spot of embodied cognition
Large Language Models (LLMs) demonstrate a kind of linguistic competence that theories of embodied and enactive cognition have long deemed impossible for systems lacking the meaningful perspective of a living being, i.e., the capacity for sense-making. Facing up to this unexpected technological development requires confronting what I propose to call the “AI dilemma”: either frontier LLMs are capable of sense-making despite lacking biological embodiment, or the kind of linguistic competence they exhibit does not necessarily require sense-making. In their chapter on cognition, Frank, Thompson, and Gleiser (2024) maintain that no AI system comes close to realizing relevance, a position that derives much of its motivation from past practical failures. However, frontier LLMs have effectively overcome Dreyfus’ commonsense knowledge problem, such that their dismissal as categorically mindless risks undermining Frank et al.’s central claim that human cognition is deeply intertwined with lived experience. I therefore argue in favor of the alternative side of the AI dilemma: human-level linguistic competence of LLMs should be recognized as a novel non‑biological form of sense‑making, based on a technologically‑mediated embodiment whose enabling properties are in need of further theoretical analysis. This reorientation invites enactive theory to clarify which aspects of sense-making may be universal and which aspects are specifically contingent on organic life, thereby advancing its conceptual framework in dialogue with contemporary AI.
AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking
The proliferation of artificial intelligence (AI) tools has transformed numerous aspects of daily life, yet its impact on critical thinking remains underexplored. This study investigates the relationship between AI tool usage and critical thinking skills, focusing on cognitive offloading as a mediating factor. Utilising a mixed-method approach, we conducted surveys and in-depth interviews with 666 participants across diverse age groups and educational backgrounds. Quantitative data were analysed using ANOVA and correlation analysis, while qualitative insights were obtained through thematic analysis of interview transcripts. The findings revealed a significant negative correlation between frequent AI tool usage and critical thinking abilities, mediated by increased cognitive offloading. Younger participants exhibited higher dependence on AI tools and lower critical thinking scores compared to older participants. Furthermore, higher educational attainment was associated with better critical thinking skills, regardless of AI usage. These results highlight the potential cognitive costs of AI tool reliance, emphasising the need for educational strategies that promote critical engagement with AI technologies. This study contributes to the growing discourse on AI’s cognitive implications, offering practical recommendations for mitigating its adverse effects on critical thinking. The findings underscore the importance of fostering critical thinking in an AI-driven world, making this research essential reading for educators, policymakers, and technologists.

AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking
The proliferation of artificial intelligence (AI) tools has transformed numerous aspects of daily life, yet its impact on critical thinking remains underexplored. This study investigates the relationship between AI tool usage and critical thinking skills, focusing on cognitive offloading as a mediating factor. Utilising a mixed-method approach, we conducted surveys and in-depth interviews with 666 participants across diverse age groups and educational backgrounds. Quantitative data were analysed using ANOVA and correlation analysis, while qualitative insights were obtained through thematic analysis of interview transcripts. The findings revealed a significant negative correlation between frequent AI tool usage and critical thinking abilities, mediated by increased cognitive offloading. Younger participants exhibited higher dependence on AI tools and lower critical thinking scores compared to older participants. Furthermore, higher educational attainment was associated with better critical thinking skills, regardless of AI usage. These results highlight the potential cognitive costs of AI tool reliance, emphasising the need for educational strategies that promote critical engagement with AI technologies. This study contributes to the growing discourse on AI’s cognitive implications, offering practical recommendations for mitigating its adverse effects on critical thinking. The findings underscore the importance of fostering critical thinking in an AI-driven world, making this research essential reading for educators, policymakers, and technologists.

We Need to Talk About How We Talk About 'AI'
We share a responsibility to create and use empowering metaphors rather than misleading language, write Emily M. Bender and Nanna Inie.

AI, Decomputing and the Interregnum
This paper treats AI as diagnostic for the deeper changes taking place in the existing order of things. It uses AI's alignment with both the political economy and with the dualisms that underpin it, including race, gender and anthropocentrism, to highlight the nihilistic character of the current restructuring. AI's scaling and accelerationism are taken as examples of the wider tactics being invoked by hegemonic power to maintain control under changing conditions. From this perspective, the massive build-out of data centres isn't simply a seizure of energy resources but a manifestation of an aggressive and misogynist technopolitics. The paper argues that a liberal push for digital sovereignty doesn't interrupt these dynamics but plays into the hands of emerging technofascism. It proposes instead the prefigurative tactic of 'decomputing', which draws on degrowth, deautomatisation and a convivial approach to technology. It explores decomputing as a means to mitigate both material and relational harms and as a decisive turn towards infrastructuring the common good. The paper concludes that AI is the contradiction that reveals many others, not least the gap between claims to legitimacy and the actuality of destructive violence, and proposes an alternative technopolitics of reciprocity that prioritises care and sustainability.
AI, Decomputing and the Interregnum
This paper treats AI as diagnostic for the deeper changes taking place in the existing order of things. It uses AI's alignment with both the political economy and with the dualisms that underpin it, including race, gender and anthropocentrism, to highlight the nihilistic character of the current restructuring. AI's scaling and accelerationism are taken as examples of the wider tactics being invoked by hegemonic power to maintain control under changing conditions. From this perspective, the massive build-out of data centres isn't simply a seizure of energy resources but a manifestation of an aggressive and misogynist technopolitics. The paper argues that a liberal push for digital sovereignty doesn't interrupt these dynamics but plays into the hands of emerging technofascism. It proposes instead the prefigurative tactic of 'decomputing', which draws on degrowth, deautomatisation and a convivial approach to technology. It explores decomputing as a means to mitigate both material and relational harms and as a decisive turn towards infrastructuring the common good. The paper concludes that AI is the contradiction that reveals many others, not least the gap between claims to legitimacy and the actuality of destructive violence, and proposes an alternative technopolitics of reciprocity that prioritises care and sustainability.
Stochastic Parrots | Emily M. Bender - Professor of Computational Linguistics
Illusions of Understanding from Outsourcing Thinking to LLMs
Some illusions of understanding are an inevitable part of the research process, while others can be avoided or overcome by careful critical thinking and observation. We are facing an increased risk of avoidable illusions as more research activities are delegated to large language models (LMM). LLMs can be useful but they cannot think, and their use can undermine our thinking and understanding. Thinking for ourselves is hard and error prone but worthwhile - and there are no shortcuts to understanding.
The homogenizing effect of large language models on human expression and thought
AbstractCognitive diversity, reflected in variations of language, perspective, and reasoning, is essential to creativity and collective intelligence. This diversity is rich and grounded in culture, history, and individual experience. Yet, as large language models (LLMs) become deeply embedded in people's lives, they risk standardizing language and reasoning. We synthesize evidence across linguistics, psychology, cognitive science, and computer science to show how LLMs reflect and reinforce dominant styles while marginalizing alternative voices and reasoning strategies. We examine how their design and widespread use contribute to this effect by mirroring patterns in their training data and amplifying convergence as all people increasingly rely on the same models across contexts. Unchecked, this homogenization risks flattening the cognitive landscapes that drive collective intelligence and adaptability.
