Large Language Models Require Curated Context for Reliable Political Fact-Checking—Even with Reasoning and Web Search
Large language models (LLMs) have raised hopes for automated end-to-end fact-checking, but prior studies report mixed results. As mainstream chatbots increasingly ship with reasoning capabilities and web search tools—and millions of users already rely on them for verification—rigorous evaluation is urgent. We evaluate 15 recent LLMs from OpenAI, Google, Meta, and DeepSeek on more than 6,000 claims fact-checked by PolitiFact, comparing standard models with reasoning- and web-search variants. Standard models perform poorly, reasoning offers minimal benefits, and web search provides only moderate gains, despite fact-checks being available on the web. In contrast, a curated RAG system using PolitiFact summaries improved macro F1 by 233% on average across model variants. These findings suggest that giving models access to curated high-quality context is a promising path for automated fact-checking.
Persuasion Index: A Theory-Guided Framework for Persuasion Analysis
Identifying persuasive rhetorical cues is critical across domains, from detecting information manipulation and improving AI safety to advancing public health communication. We propose Persuasion...

Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
Large language models (LLMs) are increasingly capable of generating personalized, persuasive text at scale, raising new questions about bias and fairness in automated communication. This paper presents the first systematic analysis of how LLMs behave when tasked with demographic-conditioned targeted messaging. We introduce a controlled evaluation framework using three leading models: GPT-4o, Llama-3.3, and Mistral-Large-2.1, across two generation settings: Standalone Generation, which isolates intrinsic demographic effects, and Context-Rich Generation, which incorporates thematic and regional context to emulate realistic targeting. We evaluate generated messages along three dimensions: lexical content, language style, and persuasive framing. We instantiate this framework on climate communication and find consistent age- and gender-based asymmetries across models: male- and youth-targeted messages tend to emphasize more assertive and progressive framing, while female- and senior-targeted messages more often reflect warmth, care, and traditional themes. Contextual prompts systematically amplify these disparities, with persuasion scores being higher for male-targeted messages, while age-related differences vary across models. Our findings demonstrate how demographic stereotypes can surface and intensify in LLM-generated targeted communication, underscoring the need for bias-aware generation pipelines and transparent auditing frameworks that explicitly account for demographic conditioning in socially sensitive applications.

Introducing AI to an Online Petition Platform Changed Outputs but not Outcomes
The rapid integration of AI writing tools into online platforms raises critical questions about their impact on content production and outcomes. We leverage a unique natural experiment on Change$.$org, a leading social advocacy platform, to causally investigate the effects of an in-platform ''write with AI'' tool. To understand the impact of the AI integration, we collected 1.5 million petitions and employed a difference-in-differences analysis. Our findings reveal that in-platform AI access significantly altered the lexical features of petitions and increased petition homogeneity, but did not improve petition outcomes. We confirmed the results in a separate analysis of repeat petition writers who wrote petitions before and after introduction of the AI tool. The results suggest that while AI writing tools can profoundly reshape online content, their practical utility for improving desired outcomes may be less beneficial than anticipated, and introduce unintended consequences like content homogenization.

Paper Skygest: Personalized Academic Recommendations on Bluesky
We build, deploy, and evaluate Paper Skygest, a custom personalized social feed for scientific content posted by a user's network on Bluesky and the AT Protocol. We leverage a new capability on emerging decentralized social media platforms: the ability for anyone to build and deploy feeds for other users, to use just as they would a native platform-built feed. To our knowledge, Paper Skygest is the first and largest such continuously deployed personalized social media feed by academics, with over 50,000 weekly uses by over 1,000 daily active users, all organically acquired. First, we quantitatively and qualitatively evaluate Paper Skygest usage, showing that it has sustained usage and satisfies users; we further show adoption of Paper Skygest increases a user's interactions with posts about research, and how interaction rates change as a function of post order. Second, we share our full code and describe our system architecture, to support other academics in building and deploying such feeds sustainably. Third, we overview the potential of custom feeds such as Paper Skygest for studying algorithm designs, building for user agency, and running recommender system experiments with organic users without partnering with a centralized platform.

Afternoon parallel: Baird Howland explores shifts in White House media in the Political Polarization and Discourse session: | IC2S2
Afternoon parallel: Baird Howland explores shifts in White House media in the Political Polarization and Discourse session:
Incorrect Citation Association for Articles in Online-Only Springer Nature Journals
We show that citation metrics of journal articles in many of the online-only Springer Nature journals and associated ones are distorted, going back to articles from 2001. We find that most likely due to an API response error, there are many incorrect references which typically lead to Article Number 1 of a given Volume. Among others, the issue affects journals such as Scientific Reports, Nature Communications, Communications journals, Cell Death & Disease, Light: Science & Applications, as well as many BMC, Discovery and npj journals. Beyond the negative effect of introducing incorrect reference information, this distorts the citation statistics of articles in these journals, with a few articles being massively over-cited compared to their peers, while many lose citations; e.g. both in Scientific Reports and in Nature Communications, 5 of the 10 top cited articles have article numbers of 1. We validate the distorted statistics by assessing data from multiple scientific literature databases: Crossref, OpenCitations, Semantic Scholar, and the journals' websites. The issue primarily arises from the inconsistent transition from page-based referencing of articles to article number-based referencing, as well as the improper handling of the change in the publisher's article metadata API. It seems that the most pressing problem has been present since approximately 2011, which we estimate affects the citation count of millions of authors.

The Due Process Deficit: Auditing AI Governance in U.S. Higher Education
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

Leader-driven or Leaderless: How Participation Structure Sustains Engagement and Shapes Narratives in Online Hate Communities
Extremist communities increasingly rely on social media to sustain and amplify divisive discourse. However, the relationship between their internal participation structures, audience engagement, and narrative expression remains underexplored. This study analyzes ten years of Facebook activity by hate groups related to the Israel–Palestine conflict, focusing on anti-Semitic and Islamophobic ideologies. Consistent with prior work, we find that higher participation centralization in online hate groups is associated with greater user engagement across hate ideologies, suggesting the role of key actors in sustaining group activity over time. Meanwhile, our narrative frame detection models—based on an eight-frame extremist taxonomy (e.g., dehumanization, violence justification)—reveal a clear contrast across hate ideologies: centralized Islamophobic groups employ more uniform messaging, while centralized anti-Semitic groups demonstrate greater framing diversity and topical breadth, potentially reflecting distinct historical trajectories and leader coordination patterns. Analysis of the inter-group network indicates that, although centralization and homophily are not clearly linked, ideological distinctions emerge: Islamophobic groups cluster tightly, whereas anti-Semitic groups remain more evenly connected. Overall, these findings clarify how participation structure may shape the dissemination pattern and resonance of extremist narratives online and provide a foundation for tailored strategies to disrupt or mitigate such discourse.
Collection
A collection of papers from IC2S2 conference
This morning at #ic2s2 I have the chance to present ongoing work on applying a SEM from survey methods to LLM text annotations: 👉 Evaluating LLM Text Annotations Without Ground Truth 📍10:45, Mansfield (210), Understanding LLMs w/ @maximiliankreutner.bsky.social Alex Cernat & @mstrohm.bsky.social
Presented our work on how friendship muda lets the expression of hierarchy in cooperation for elementary school children on the lighting talks at #ic2s2 #ic2s22026 @ic2s2.bsky.social today!
Attending my first #IC2S2 at Burlington, VT! I’ll be presenting my poster on July 30 (Thursday), 1:15-2:15 pm at Davis Center @uvmvermont.bsky.social If you are attending @ic2s2.bsky.social & interested in bias auditing, please stop by. Thanks @vcsi.bsky.social for sponsoring registration #ic2s22026
At today's #IC2S2 Methods and Measurement session, Linnea Gandhi @linneagandhi.bsky.social presented her study with Elizabeth Tipton and Duncan Watts, which tests whether research syntheses can predict the results of new studies and highlights the challenges of producing generalizable knowledge.
#ic2s2 opening slides
I’ll be at #IC2S2 this week! I will be presenting two papers from my lab, led by my amazing students @laylab.bsky.social & @elhamagk.bsky.social. If you’re interested in narratives,mental health, stigma, or NLP, let’s connect! Looking forward to catching up with old friends, and meeting new people.
Wonderful talk by @sjgreenwood.bsky.social about @paper-feed.bsky.social! Custom feeds on Bluesky are unique resources for researchers interested in recommendation algorithms #IC2S2 Featuring many Bluesky friends like @graze.social @devingaffney.com @kissane.myatproto.social @aendra.com
#IC2S2 keynote Gasper Begus giving a fascinating talk about analyzing languages spoken by whales @ic2s2.bsky.social
Lucas Gautheron conducts a really interesting online design experiment to show popularity feedback kills diversity and innovation. Super cool work!! #IC2S2 @ic2s2.bsky.social
Vicky Yang studies when social influence helps/hurts collective binary decision making #IC2S2 @ic2s2.bsky.social