







Replication dataset for "State Media Control Influences Large Language Models," forthcoming in Nature (https://doi.org/10.1038/s41586-026-10506-7). We show through six studies that government control of the media across the world influences the output of large language models (LLMs) via their training data.
Replication Data for "State Media Control Influences Large Language Models"
Replication dataset for "State Media Control Influences Large Language Models," forthcoming in Nature (https://doi.org/10.1038/s41586-026-10506-7). We show through six studies that government control of the media across the world influences the output of large language models (LLMs) via their training data.
State Media Control Influences Large Language Models – State Media & LLMs
Hannah Waight1,2, Eddie Yang1,3, Yin Yuan4, Solomon Messing5, Margaret E. Roberts4, Brandon M. Stewart6, Joshua A. Tucker5,7
State media control influences large language models
Millions of people around the world query large language models (LLMs) for information. Although several studies have compellingly documented the persuasive potential of these models1–10, there is limited evidence of who or what influences the models themselves, leading to a flurry of concerns about which companies and governments build and regulate the models. Here we show through six studies that government control of the media across the world already influences the output of LLMs via their training data. We use a cross-national audit to show that LLMs exhibit a stronger pro-government valence in the languages of countries with lower media freedom than in those with higher media freedom. This result is correlational, so to triangulate the specific mechanism of how state media control can influence LLMs, we develop a multi-part case study on China’s media. We demonstrate that media scripted and curated by the Chinese state appears in LLM training datasets. To evaluate the plausible effect of this inclusion, we use an open-weight model to show that additional pretraining on Chinese state-coordinated media generates more positive answers to prompts about Chinese political institutions and leaders. We link this phenomenon to commercial models through two audit studies demonstrating that prompting models in Chinese generates more positive responses about China’s institutions and leaders than do the same queries in English. The combination of influence and persuasive potential across languages suggests the troubling conclusion that states and powerful institutions have increased strategic incentives to leverage media control in the hopes of shaping LLM output.

State media control influences large language models
Millions of people around the world query large language models (LLMs) for information. Although several studies have compellingly documented the persuasive potential of these models1–10, there is limited evidence of who or what influences the models themselves, leading to a flurry of concerns about which companies and governments build and regulate the models. Here we show through six studies that government control of the media across the world already influences the output of LLMs via their training data. We use a cross-national audit to show that LLMs exhibit a stronger pro-government valence in the languages of countries with lower media freedom than in those with higher media freedom. This result is correlational, so to triangulate the specific mechanism of how state media control can influence LLMs, we develop a multi-part case study on China’s media. We demonstrate that media scripted and curated by the Chinese state appears in LLM training datasets. To evaluate the plausible effect of this inclusion, we use an open-weight model to show that additional pretraining on Chinese state-coordinated media generates more positive answers to prompts about Chinese political institutions and leaders. We link this phenomenon to commercial models through two audit studies demonstrating that prompting models in Chinese generates more positive responses about China’s institutions and leaders than do the same queries in English. The combination of influence and persuasive potential across languages suggests the troubling conclusion that states and powerful institutions have increased strategic incentives to leverage media control in the hopes of shaping LLM output.

State Media Monitor Global Dataset 2025
This dataset provides comprehensive information on the governance, funding, and editorial independence of state and public media outlets worldwide. In 2025, the State Media Monitor covers 170 countries, tracking changes in media governance models, funding mechanisms, and degrees of political control. The dataset underpins the annual State Media Monitor Global Study and supports comparative research in journalism, media policy, and governance. Access the full dataset here: https://www.statemediamonitor.com
How Large Language Models Actually Work
Language Models Trained on State Media Sources Launder Propaganda
Tech Policy Press is a nonprofit media and community venture intended to provoke new ideas, debate and discussion at the intersection of technology and democracy. We publish opinion and analysis.

LLMs and World Models, Part 1
How do Large Language Models Make Sense of Their “Worlds”?

Communication Bias in Large Language Models: A Regulatory Perspective
Large language models (LLMs) are increasingly central to many applications, raising concerns about bias, fairness, and regulatory compliance. This paper reviews risks of biased outputs and their...

Large language model
A large language model (LLM) is a neural network trained on a vast amount of text for natural language processing tasks, especially language generation. LLMs can typically generate, summarize, translate, and analyze text in many contexts, and are a foundational technology behind modern chatbots.[1] Biased or inaccurate training data can make an LLM's output less reliable.[2]
Large language models reduce public knowledge sharing on online Q&A platforms
Abstract. Large language models (LLMs) are a potential substitute for human-generated data and knowledge resources. This substitution, however, can present

Scaling Laws Across Model Architectures: A Comparative Analysis of...
The scaling of large language models (LLMs) is a critical research area for the efficiency and effectiveness of model training and deployment. Our work investigates the transferability and...

AI Large Language Model Training: The Potential Risks of Ideological Skewing — PSG Consulting
LLMs (AI Large Language Models) have become part of everyday life. Systems such as ChatGPT, Claude, Gemini, Meta AI (Llama) and X.ai's Grok handle billions of interactions daily. They increasingly shape what information people encounter and in what order, subtly deciding what's important and even what is true, sometimes without users realizing it. Because LLMs wield growing power over information exposure, it is vital to recognize the political and ideological structures at multiple stages of their design, and to identify manipulation risks.

Large language models are cultural technologies. What might that mean?
Four different perspectives

Take caution in using LLMs as human surrogates | PNAS
Recent studies suggest large language models (LLMs) can generate human-like responses, aligning with human behavior in economic experiments, survey...

Are LLMs Stifling Political Speech? An Assessment of How AI Models Protect Free Expression | Oversight Board
The Oversight Board’s first evaluation of large language models (LLMs) shows that some of the world’s most-used models from Anthropic, DeepSeek, Google, Meta