







Fully automatic censorship removal for language models - p-e-w/heretic
Robust Steganography from Large Language Models
Recent steganographic schemes, starting with Meteor (CCS'21), rely on leveraging large language models (LLMs) to resolve a historically-challenging task of disguising covert communication as ``innocent-looking'' natural-language communication. However, existing methods are vulnerable to ``re-randomization attacks,'' where slight changes to the communicated text, that might go unnoticed, completely destroy any hidden message. This is also a vulnerability in more traditional encryption-based stegosystems, where adversaries can modify the randomness of an encryption scheme to destroy the hidden message while preserving an acceptable covertext to ordinary users. In this work, we study the problem of robust steganography. We introduce formal definitions of weak and strong robust LLM-based steganography, corresponding to two threat models in which natural language serves as a covertext channel resistant to realistic re-randomization attacks. We then propose two constructions satisfying these notions. We design and implement our steganographic schemes that embed arbitrary secret messages into natural language text generated by LLMs, ensuring recoverability even under adversarial paraphrasing and rewording attacks. To support further research and real-world deployment, we release our implementation and datasets for public use.

The Consent Layer: Using ligatures to make web text expensive to scrape without asking
ShieldFont is an open-source creative technology project that offers a practical opt-out from unauthorized AI training and disrupts what is collected when that choice is ignored. It swaps 45.8% of content words (around 24.4% of all words) in a page's source code for other (partially) random words, while the font restores the original text on screen. Readers see the work as intended; mass scrapers collect an altered version. In testing, shielding caused over 90% of pages that would otherwise pass the quality filter to be rejected, keeping them out of the training pipeline. Of those that still passed, 19.4% of all words conveyed false meaning, adding noise to unauthorized AI training datasets. This paper's goal is to walk newcomers through the whole process, in plain language and in order: the project's rationale, how it was built, the results, how to deploy it, and where to contribute.
Are LLMs Stifling Political Speech? An Assessment of How AI Models Protect Free Expression | Oversight Board
The Oversight Board’s first evaluation of large language models (LLMs) shows that some of the world’s most-used models from Anthropic, DeepSeek, Google, Meta
Removing sexually explicit content from r/all
327K subscribers in the modnews community. An official community for announcements from Reddit, Inc. pertaining to moderation.

Forcing Generative Models to Degenerate Ones: The Power of Data...
Growing applications of large language models (LLMs) trained by a third party raise serious concerns on the security vulnerability of LLMs.It has been demonstrated that malicious actors can...

Wikipedia Bans AI-Generated Content
“In recent months, more and more administrative reports centered on LLM-related issues, and editors were being overwhelmed.”
Jailbreaking Large Language Models: If You Torture the Model Long Enough, It Will Confess!
A Cautionary Tale…

Meta: Systemic Censorship of Palestine Content
Meta’s content moderation policies and systems have increasingly silenced voices in support of Palestine on Instagram and Facebook in the wake of the hostilities between Israeli forces and Palestinian armed groups.

Political Neutrality in AI Is Impossible- But Here Is How to Approximate It
AI systems often exhibit political bias, influencing users' opinions and decisions. While political neutrality-defined as the absence of bias-is often seen as an ideal solution for fairness and safety, this position paper argues that true political neutrality is neither feasible nor universally desirable due to its subjective nature and the biases inherent in AI training data, algorithms, and user interactions. However, inspired by Joseph Raz's philosophical insight that "neutrality [...] can be a matter of degree" (Raz, 1986), we argue that striving for some neutrality remains essential for promoting balanced AI interactions and mitigating user manipulation. Therefore, we use the term "approximation" of political neutrality to shift the focus from unattainable absolutes to achievable, practical proxies. We propose eight techniques for approximating neutrality across three levels of conceptualizing AI, examining their trade-offs and implementation strategies. In addition, we explore two concrete applications of these approximations to illustrate their practicality. Finally, we assess our framework on current large language models (LLMs) at the output level, providing a demonstration of how it can be evaluated. This work seeks to advance nuanced discussions of political neutrality in AI and promote the development of responsible, aligned language models.

Political Neutrality in AI Is Impossible- But Here Is How to Approximate It
AI systems often exhibit political bias, influencing users' opinions and decisions. While political neutrality-defined as the absence of bias-is often seen as an ideal solution for fairness and safety, this position paper argues that true political neutrality is neither feasible nor universally desirable due to its subjective nature and the biases inherent in AI training data, algorithms, and user interactions. However, inspired by Joseph Raz's philosophical insight that "neutrality [...] can be a matter of degree" (Raz, 1986), we argue that striving for some neutrality remains essential for promoting balanced AI interactions and mitigating user manipulation. Therefore, we use the term "approximation" of political neutrality to shift the focus from unattainable absolutes to achievable, practical proxies. We propose eight techniques for approximating neutrality across three levels of conceptualizing AI, examining their trade-offs and implementation strategies. In addition, we explore two concrete applications of these approximations to illustrate their practicality. Finally, we assess our framework on current large language models (LLMs) at the output level, providing a demonstration of how it can be evaluated. This work seeks to advance nuanced discussions of political neutrality in AI and promote the development of responsible, aligned language models.

Moderation on Blacksky-Only Posts: A Community Proposal
Our team has explored several ways to make permissioned posts possible. Now we’re bringing the decision back to the community. Below are three approaches to moderation: a machine-learning system that automates moderation, a large-scale peer moderator program, or a hybrid model that combines both. We’re inviting your feedback to help decide how we should govern privacy together.

Online Safety Bills Are Fueling a New Wave of Internet Censorship
State and federal bills seek to limit minors’ access to social media, but civil liberties advocates warn that the resulting online censorship threatens constitutional rights without delivering real safety.

Recursive Language Models: the paradigm of 2026
How we plan to manage extremely long contexts
.png?v=c8c07d4bf43b)
I wish the Atmosphere had content warnings (like on the fediverse)... because I'd like to show you the limits of free speech absolutism and the kind of content that is allowed on #WSocial: blatant racism, xenophobia and transphobia. X vibes... but hey, it's "hosted in Europe" 😒
The forced global adoption of puritanistic American cultural norms vis sexuality and nudity via app store/payment processor rules is one of the least-discussed free speech issues of the last decade and this lack of attention is going to hurt online expression so much in the years to come.
♍ Bee 🐝
I think also we cannot ignore the potential social and cultural ramifications of normalizing the demonification of anything adult or sexual The whole system is stacked against any sort of movement in the other direction and it's not something the porn industry can handle alone at this point