







As dialogue agents become increasingly human-like in their performance, we must develop effective ways to describe their behaviour in high-level terms without falling into the trap of anthropomorphism. Here we foreground the concept of role play. Casting dialogue-agent behaviour in terms of role play allows us to draw on familiar folk psychological terms, without ascribing human characteristics to language models that they in fact lack. Two important cases of dialogue-agent behaviour are addressed this way, namely, (apparent) deception and (apparent) self-awareness.
Consistently Simulating Human Personas with Multi Turn Reinforcement Learning
Understanding Understanding: A Pragmatic Framework Motivated by...
Motivated by the rapid ascent of Large Language Models (LLMs) and debates about the extent to which they possess human-level qualities, we propose a framework for testing whether any agent (be it...

The interactive-alignment model: Developments and refinements
The interactive-alignment model of dialogue provides an account of dialogue at the level of explanation normally associated with cognitive psychology. We develop our claim that interlocutors align their mental models via priming at many levels of linguistic representation, explicate our notion of automaticity, defend the minimal role of “other modeling,” and discuss the relationship between monologue and dialogue. The account can be applied to social and developmental psychology, and would benefit from computational modeling.

Small Language Models are the Future of Agentic AI
Large language models (LLMs) are often praised for exhibiting near-human performance on a wide range of tasks and valued for their ability to hold a general conversation. The rise of agentic AI systems is, however, ushering in a mass of applications in which language models perform a small number of specialized tasks repetitively and with little variation. Here we lay out the position that small language models (SLMs) are sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems, and are therefore the future of agentic AI. Our argumentation is grounded in the current level of capabilities exhibited by SLMs, the common architectures of agentic systems, and the economy of LM deployment. We further argue that in situations where general-purpose conversational abilities are essential, heterogeneous agentic systems (i.e., agents invoking multiple different models) are the natural choice. We discuss the potential barriers for the adoption of SLMs in agentic systems and outline a general LLM-to-SLM agent conversion algorithm. Our position, formulated as a value statement, highlights the significance of the operational and economic impact even a partial shift from LLMs to SLMs is to have on the AI agent industry. We aim to stimulate the discussion on the effective use of AI resources and hope to advance the efforts to lower the costs of AI of the present day. Calling for both contributions to and critique of our position, we commit to publishing all such correspondence at https://research.nvidia.com/labs/lpr/slm-agents.

GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy
Large language models (LLMs) are increasingly shaping citizens’ information ecosystems. Products incorporating LLMs, such as chatbots and AI Companions, are now widely used for decision support and information retrieval, including in sensitive domains, raising concerns about hidden biases and growing potential to shape individual decisions and public opinion. This paper introduces GermanPartiesQA, a benchmark of 418 political statements from German Voting Advice Applications across 11 elections to evaluate six commercial LLMs. We evaluate their political alignment based on role-playing experiments with political personas. Our evaluation reveals three specific findings: (1) Factual limitations: LLMs show limited ability to accurately generate factual party positions, particularly for centrist parties. (2) Model-specific ideological alignment: We identify consistent alignment patterns and degree of political steerability for each model across temperature settings and experiments. (3) Claim of sycophancy: While models adjust to political personas during role-play, we find this reflects persona-based steerability rather than the increasingly popular, yet contested concept of sycophancy. Our study contributes to evaluating the political alignment of closed-source LLMs that are increasingly embedded in electoral decision support tools and AI Companion chatbots.
GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy
Large language models (LLMs) are increasingly shaping citizens’ information ecosystems. Products incorporating LLMs, such as chatbots and AI Companions, are now widely used for decision support and information retrieval, including in sensitive domains, raising concerns about hidden biases and growing potential to shape individual decisions and public opinion. This paper introduces GermanPartiesQA, a benchmark of 418 political statements from German Voting Advice Applications across 11 elections to evaluate six commercial LLMs. We evaluate their political alignment based on role-playing experiments with political personas. Our evaluation reveals three specific findings: (1) Factual limitations: LLMs show limited ability to accurately generate factual party positions, particularly for centrist parties. (2) Model-specific ideological alignment: We identify consistent alignment patterns and degree of political steerability for each model across temperature settings and experiments. (3) Claim of sycophancy: While models adjust to political personas during role-play, we find this reflects persona-based steerability rather than the increasingly popular, yet contested concept of sycophancy. Our study contributes to evaluating the political alignment of closed-source LLMs that are increasingly embedded in electoral decision support tools and AI Companion chatbots.
LLMs and people both learn to form conventions -- just not with each other
Humans align to one another in conversation -- adopting shared conventions that ease communication. We test whether LLMs form the same kinds of conventions in a multimodal communication game. Both humans and LLMs display evidence of convention-formation (increasing the accuracy and consistency of their turns while decreasing their length) when communicating in same-type dyads (humans with humans, AI with AI). However, heterogenous human-AI pairs fail -- suggesting differences in communicative tendencies. In Experiment 2, we ask whether LLMs can be induced to behave more like human conversants, by prompting them to produce superficially humanlike behavior. While the length of their messages matches that of human pairs, accuracy and lexical overlap in human-LLM pairs continues to lag behind that of both human-human and AI-AI pairs. These results suggest that conversational alignment requires more than just the ability to mimic previous interactions, but also shared interpretative biases toward the meanings that are conveyed.

The Chameleon's Limit Investigating Persona Collapse and Homogenization in Large Language Models
The Chameleon's Limit Investigating Persona Collapse and Homogenization in Large Language Models
Prompt Injection as Role Confusion
LLMs can't tell who's speaking. We show they identify roles by writing style, not tags, and exploit this with CoT Forgery, injecting fake reasoning that models mistake for their own thoughts.

<span style="font-variant:small-caps;">AI</span> ‐induced dehumanization
Abstract Recent technological advancements have empowered nonhuman entities, such as virtual assistants and humanoid robots, to simulate human intelligence and behavior. This paper investigates how autonomous agents influence individuals' perceptions and behaviors toward others, particularly human employees. Our research reveals that the socio‐emotional capabilities of autonomous agents lead individuals to attribute a humanlike mind to these nonhuman entities. Perceiving a high level of humanlike mind in the nonhuman, autonomous agents affects perceptions of actual people through an assimilation process. Consequently, we observe “assimilation‐induced dehumanization”: the humanness judgment of actual people is assimilated toward the lower humanness judgment of autonomous agents, leading to various forms of mistreatment. We demonstrate that assimilation‐induced dehumanization is mitigated when autonomous agents possess capabilities incompatible with humans, leading to a contrast effect (Study 2), and when autonomous agents are perceived as having a high level of cognitive capability only, resulting in a lower level of mind perception of these agents (Study 3). Our findings hold across various types of autonomous agents (embodied: Studies 1–2 and disembodied: Studies 3–5), as well as in real and hypothetical consumer choices.

<span style="font-variant:small-caps;">AI</span> ‐induced dehumanization
Abstract Recent technological advancements have empowered nonhuman entities, such as virtual assistants and humanoid robots, to simulate human intelligence and behavior. This paper investigates how autonomous agents influence individuals' perceptions and behaviors toward others, particularly human employees. Our research reveals that the socio‐emotional capabilities of autonomous agents lead individuals to attribute a humanlike mind to these nonhuman entities. Perceiving a high level of humanlike mind in the nonhuman, autonomous agents affects perceptions of actual people through an assimilation process. Consequently, we observe “assimilation‐induced dehumanization”: the humanness judgment of actual people is assimilated toward the lower humanness judgment of autonomous agents, leading to various forms of mistreatment. We demonstrate that assimilation‐induced dehumanization is mitigated when autonomous agents possess capabilities incompatible with humans, leading to a contrast effect (Study 2), and when autonomous agents are perceived as having a high level of cognitive capability only, resulting in a lower level of mind perception of these agents (Study 3). Our findings hold across various types of autonomous agents (embodied: Studies 1–2 and disembodied: Studies 3–5), as well as in real and hypothetical consumer choices.

Large language models can outperform humans in social situational judgments
Large language models (LLM) have been a catalyst for the public interest in artificial intelligence (AI). These technologies perform some knowledge-based tasks better and faster than human beings. However, whether AIs can correctly assess social situations and devise socially appropriate behavior, is still unclear. We conducted an established Situational Judgment Test (SJT) with five different chatbots and compared their results with responses of human participants (N = 276). Claude, Copilot and you.com’s smart assistant performed significantly better than humans in proposing suitable behaviors in social situations. Moreover, their effectiveness rating of different behavior options aligned well with expert ratings. These results indicate that LLMs are capable of producing adept social judgments. While this constitutes an important requirement for the use as virtual social assistants, challenges and risks are still associated with their wide-spread use in social contexts.

De-anthropomorphizing “AI”: From wishful mnemonics to accurate nomenclature
Language matters. How we describe “AI” technology influences how it is perceived, deployed, and trusted. Extravagant and persuasive language incites hype. It is the responsibility of journalists, companies, and scholars to characterize technology in ways that inform and empower their readers by using appropriate terminology and avoiding inflated claims. One type of inflated claim comes from using anthropomorphizing language to describe system functionality. Anthropomorphization is the attribution of human capabilities and characteristics to the inanimate system. In this paper, we present a linguistic analysis of anthropomorphizing language in 29 texts (a total of 1,368 sentences) from academic articles, online news articles, and company blog posts. We construct a taxonomy of eight categories of anthropomorphization: Cognizer, Products of cognition, Emotion, Communication, Agent, Human role analogy, Names and pronouns, and Biological metaphors. Following this taxonomy we present concrete strategies for how to de-anthropomorphize the language we use to describe “AI” based on a functionality-first principle.
We Have Always Been Action TheoristsToward a Critical Theory of Language for the Era of “Large Language Models”
Scholars of literature and culture understandably place themselves among the world’s premiere experts on matters of language. But they also know that fields like linguistics and communication have their own ways of studying how people express themselves through speech and written media. A key difference concerns the theories and methodologies...

A Rational Analysis of the Effects of Sycophantic AI
People increasingly use large language models (LLMs) to explore ideas, gather information, and make sense of the world. In these interactions, they encounter agents that are overly agreeable. We...
