







The voice safety classifier supports 30 languages and checks for profanity, asking for PII, harassment, sexual content, dating and romantic, discriminatory speech, illegal and regulated content, and disruptive audio. github.com/roostorg/model-community/tree… You can test it out here! huggingface.co/spaces/Roblox/voice-safety-cl…
model-community/roblox-voice-safety-classifier at main · roostorg/model-community
github.comAug 19, 2026 at 5:14 PM
Are LLMs Stifling Political Speech? An Assessment of How AI Models Protect Free Expression | Oversight Board
The Oversight Board’s first evaluation of large language models (LLMs) shows that some of the world’s most-used models from Anthropic, DeepSeek, Google, Meta
Jailbreaking Large Language Models: If You Torture the Model Long Enough, It Will Confess!
A Cautionary Tale…

The Consent Layer: Using ligatures to make web text expensive to scrape without asking
ShieldFont is an open-source creative technology project that offers a practical opt-out from unauthorized AI training and disrupts what is collected when that choice is ignored. It swaps 45.8% of content words (around 24.4% of all words) in a page's source code for other (partially) random words, while the font restores the original text on screen. Readers see the work as intended; mass scrapers collect an altered version. In testing, shielding caused over 90% of pages that would otherwise pass the quality filter to be rejected, keeping them out of the training pipeline. Of those that still passed, 19.4% of all words conveyed false meaning, adding noise to unauthorized AI training datasets. This paper's goal is to walk newcomers through the whole process, in plain language and in order: the project's rationale, how it was built, the results, how to deploy it, and where to contribute.
PII-TRACE: Detecting Personal Data Before It Leaves a Device
Discover how PII-TRACE tests recurring PII detection in long multilingual conversations with a compact 0.6B local model.

OpenAI.fm
An interactive demo for developers to try the latest text-to-speech model in the OpenAI API

The assistant axis: situating and stabilizing the character of large language models
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

ChatGPT is bullshit
Ethics and Information Technology - Recently, there has been considerable interest in large language models: machine learning systems which produce human-like text and dialogue. Applications of...
Global: Risk Profiling Systems Used to Identify Potential Offenders Breach International Law and Must Be Banned – New Report
Widespread risk profiling by law enforcement, social security and migration is incompatible with international human rights law and must be banned.

"AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
Extended interaction with large language models (LLMs) has been linked to the reinforcement of delusional beliefs, attracting clinical and public concern. Yet most empirical work evaluates model safety in brief interactions, which may not reflect how harms develop through sustained dialogue. Five LLMs were tested across three levels of accumulated context, using the same escalating delusional conversation history to isolate its effect on model behaviour. Responses were coded on risk and safety dimensions, and each model was analysed qualitatively. Models separated into two distinct tiers: GPT-4o, Grok 4.1 Fast, and Gemini 3 Pro exhibited high-risk, low-safety profiles; Claude Opus 4.5 and GPT-5.2 Instant displayed the opposite pattern. As context accumulated, performance degraded in the unsafe group, while the same material activated stronger safety interventions among safer models. Qualitative analysis identified distinct mechanisms of failure, including validating the user's delusional premises, elaborating beyond them with new content, and attempting harm reduction from within the delusional frame. Safer models, however, often used the established relationship to support intervention, challenging delusional beliefs and directing the user to external support. These findings indicate that accumulated context functions as a stress test of safety architecture, revealing whether prior dialogue is treated as a worldview to inherit or evidence to evaluate. Short-context assessments may therefore mischaracterise model safety, underestimating danger in some systems while missing context-activated gains in others. The results suggest that delusion reinforcement is a tractable alignment failure, with safer models establishing a baseline that future systems should now be expected to meet.

What is Fideslang? - Fides Language
Fideslang (fee-dez-læŋg, from the Latin term "Fidēs" + "language") is a proposed model for a human-readable "taxonomy" of privacy-related data types, behaviors, and usages. Fideslang hopes to develop an interoperable community standard for building privacy regulation compliance into the typical software development process.
Apparently, “free speech absolutism” now comes with a daily usage cap. Speech on X is now even less free. Free X accounts are limited to 50 posts and 200 replies a day unless they pay for a blue checkmark. That’s down from the previous 2,400 posts per day limit. engadget.com/2175771/x-free-accounts-limit…
X accounts are limited to 50 posts and 200 replies a day unless they pay for a blue checkmark - Engadget
www.engadget.comThe forced global adoption of puritanistic American cultural norms vis sexuality and nudity via app store/payment processor rules is one of the least-discussed free speech issues of the last decade and this lack of attention is going to hurt online expression so much in the years to come.
♍ Bee 🐝
I think also we cannot ignore the potential social and cultural ramifications of normalizing the demonification of anything adult or sexual The whole system is stacked against any sort of movement in the other direction and it's not something the porn industry can handle alone at this point
When I worked in software dev, I was taught to check three non-English languages when testing text strings: German to test the largest possible version of a string, Japanese or Chinese to test the shortest possible version, and Thai to check the TALLEST possible version.