







A strategic imperative for building a subculture of innovation
The Safety Levers
A framework to move your subculture from the Anxiety Zone to the Learning Zone

Trust & Safety Tycoon
Manage your team, set policies, make investments, and tackle the challenging world of Trust & Safety

An End-to-End View of AI Safety
In this blog, Rachel Coldicutt OBE, Executive Director, Careful Industries discusses our newly published literature review on the safe adoption of artificial intelligence in engineered systems.

A Safe Path to Open Weights
Strategic openness can strengthen AI safety and support broader access as defenses and safety science mature.

(PDF) Leadership Failure in the Eyes of Subordinates: Perception, Antecedents, and Consequences *
PDF | Effective leaders and leadership behaviors and practices are well-researched areas in the domain of leadership and management. However,... | Find, read and cite all the research you need on ResearchGate

Autonomy and Clarity in Leadership Styles
There’s no one right style of leadership. Use the right style of at the right time.

"AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
Extended interaction with large language models (LLMs) has been linked to the reinforcement of delusional beliefs, attracting clinical and public concern. Yet most empirical work evaluates model safety in brief interactions, which may not reflect how harms develop through sustained dialogue. Five LLMs were tested across three levels of accumulated context, using the same escalating delusional conversation history to isolate its effect on model behaviour. Responses were coded on risk and safety dimensions, and each model was analysed qualitatively. Models separated into two distinct tiers: GPT-4o, Grok 4.1 Fast, and Gemini 3 Pro exhibited high-risk, low-safety profiles; Claude Opus 4.5 and GPT-5.2 Instant displayed the opposite pattern. As context accumulated, performance degraded in the unsafe group, while the same material activated stronger safety interventions among safer models. Qualitative analysis identified distinct mechanisms of failure, including validating the user's delusional premises, elaborating beyond them with new content, and attempting harm reduction from within the delusional frame. Safer models, however, often used the established relationship to support intervention, challenging delusional beliefs and directing the user to external support. These findings indicate that accumulated context functions as a stress test of safety architecture, revealing whether prior dialogue is treated as a worldview to inherit or evidence to evaluate. Short-context assessments may therefore mischaracterise model safety, underestimating danger in some systems while missing context-activated gains in others. The results suggest that delusion reinforcement is a tractable alignment failure, with safer models establishing a baseline that future systems should now be expected to meet.

Why Good Leaders Fail
The risk of sudden leadership failure can be headed off by early detection of challenges and better supports.

Trust & Safety Library - Trust & Safety Professional Association
Welcome to the Trust & Safety Library! We use this space to collect articles, blog posts, journal articles, lectures, podcasts, and websites that trust and safety professionals may find useful in developing policies, supporting moderators, building systems to detect violations, and generally deepening their practice. We welcome your submissions and feedback. This project was initially

Why Psychology Hasn’t Had a Big New Idea in Decades
Can our field get its act together?

Open by Design: ROOST's Approach to Safety Tool Development

Trust & Safety Curriculum - Trust & Safety Professional Association
The Trust & Safety Curriculum defines core concepts, terms, & standard practices of body of knowledge we call “trust and safety.”

consilienceproject.org/development-in-progress/ This essay argues our idea of progress is immature, ignoring the scale of harmful side effects. Using cases like leaded gasoline and synthetic fertilizer, it calls for a mature approach that internalizes externalities for a sustainable future.- LLM gen summary
Development in Progress - The Consilience Project
consilienceproject.orgWhy does every platform rebuild the same safety tools from scratch, behind closed doors? Our Head of Product @julietshen.bsky.social joined the Won't Fix pod to talk open-source T&S infrastructure, AI vs. human judgment & safety for the decentralized web: youtube.com/watch?v=RxFV7VwxkLs
Won't Fix Episode 9: With Juliet Shen, Cofounder & HOP at ROOST
www.youtube.com