







In this blog, Rachel Coldicutt OBE, Executive Director, Careful Industries discusses our newly published literature review on the safe adoption of artificial intelligence in engineered systems.
AI Researchers On AI Risk
I first became interested in AI risk back around 2007. At the time, most people’s response to the topic was “Haha, come back when anyone believes this besides random Internet crackpots.…

Labor market impacts of AI: A new measure and early evidence
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Rethinking security for the age of AI - The Official Microsoft Blog
Editor’s note: Updates with additional details on the model’s crash score. Why security needs a new cyber stack — Introducing Project Perception The physics of cybersecurity are changing. Autonomous systems can now reason, adapt and operate continuously. At the same time, the cost of offense is falling, while the volume, velocity and complexity of what...

A Safe Path to Open Weights
Strategic openness can strengthen AI safety and support broader access as defenses and safety science mature.

Human-Centered Artificial Intelligence: Three Fresh Ideas
Human-Centered AI (HCAI) is a promising direction for designing AI systems that support human self-efficacy, promote creativity, clarify responsibility, and facilitate social participation. These human aspirations also encourage consideration of privacy, security, environmental protection, social justice, and human rights. This commentary reverses the current emphasis on algorithms and AI methods, by putting humans at the center of systems design thinking, in effect, a second Copernican Revolution. It offers three ideas: (1) a two-dimensional HCAI framework, which shows how it is possible to have both high levels of human control AND high levels of automation, (2) a shift from emulating humans to empowering people with a plea to shift language, imagery, and metaphors away from portrayals of intelligent autonomous teammates towards descriptions of powerful tool-like appliances and tele-operated devices, and (3) a three-level governance structure that describes how software engineering teams can develop more reliable systems, how managers can emphasize a safety culture across an organization, and how industry-wide certification can promote trustworthy HCAI systems. These ideas will be challenged by some, refined by others, extended to accommodate new technologies, and validated with quantitative and qualitative research. They offer a reframe -- a chance to restart design discussions for products and services -- which could bring greater benefits to individuals, families, communities, businesses, and society.
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.

As the Federal Government Rushes Toward AI, Here Are Three Cautionary Tales — ProPublica
We’ve been reporting on cybersecurity for years. As President Donald Trump and his Cabinet say artificial intelligence will transform the nation, the messaging isn’t new. It follows a familiar pattern.

Developing Enterprise Frontier Safeguards with our customers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
AI labs’ all-or-nothing race leaves no time to fuss about safety
They have ideas about how to restrain wayward models, but worry that doing so will disadvantage them

Curriculum | AI Safety — ARENA
Explore the ARENA curriculum for AI safety education, covering fundamentals, transformers, reinforcement learning, and evaluations, with resources for educators and learners.
AI agents pose untold risk to humanity. We must act to prevent that future | David Krueger
The pieces are falling into place for autonomous artificial intelligence. We must stop unregulated development

Anthropic Drops Flagship Safety Pledge
In an abrupt shift, the company may release future AI models without ironclad safety guarantees

70 years of AI hype
Quoting from Olivia Guest et al. (2025) "Against the Uncritical Adoption of AI Technologies in Academia."

AI #163: Mythos Quest
There exists an AI model, Claude Mythos, that has discovered critical safety vulnerabilities in every major operating system and browser.

AI Safety Is a Narrative Problem · Special Issue 5: Grappling With the Generative AI Revolution
This op-ed explores power and narrative dynamics around AI. Drawing on pop-culture references, the professional experiences of the author and examples from 2023’s “Great AI Safety Hype Roadshow,” this piece draws on the literary criticism technique of practical criticism to consider how speeches and announcements from both Silicon Valley executives and research scientists to interrogate the media-friendly nature of p(doom) discourse—which focuses on the existential risks of AI (PauseAI, 2023)—and its likely consequences. The complexities of AI and its numerous social impacts can be difficult for even the most expert analyst to unpack. In spite of this, the potential of “existential threats” has successfully cut through to become a mainstay of mainstream media coverage over the last year. This piece will make the case that this is an effective narrative conceit that has achieved a number of ends that traditional science communication tends to find difficult, if not impossible, to achieve. Firstly, it is easy to understand. Simplification of this nature—that removes jargon and complexity and focuses on a single outcome—is much easier to fit on a TV rolling news ticker or on the cover of a tabloid newspaper than more well-balanced, representative opinions. Secondly, it inherits prior assumptions from well-known dramatic forms. P(doom) plays to stories familiar from Greek tragedy through to Marvel movies, in which lone male heroes battle ineluctable forces. Thirdly, it is imbued with urgency and so becomes difficult to ignore.
