An End-to-End View of AI Safety
In this blog, Rachel Coldicutt OBE, Executive Director, Careful Industries discusses our newly published literature review on the safe adoption of artificial intelligence in engineered systems.

AI for Science & Safety Nodes - Request for Proposals
Artificial intelligence is accelerating the pace of discovery across science and technology. But today’s AI ecosystem risks centralizing compute, talent, and decision-making power – concentrating capabilities in ways that could undermine both innovation and safety.

Rethinking security for the age of AI - The Official Microsoft Blog
Editor’s note: Updates with additional details on the model’s crash score. Why security needs a new cyber stack — Introducing Project Perception The physics of cybersecurity are changing. Autonomous systems can now reason, adapt and operate continuously. At the same time, the cost of offense is falling, while the volume, velocity and complexity of what...

A Safe Path to Open Weights
Strategic openness can strengthen AI safety and support broader access as defenses and safety science mature.

The AI-as-Normal-Technology view of loss of control incidents
A middle ground between the cybersecurity and AI safety communities

The lethal trifecta for AI agents: private data, untrusted content, and external communication
If you are a user of LLM systems that use tools (you can call them “AI agents” if you like) it is critically important that you understand the risk of …

The agent control plane gets real - Sensemaker
Two prompt-injection incidents show why agent security is about permission boundaries, not better instructions.
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.

Detecting and preventing distillation attacks
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Where agents meet the gate - Sensemaker
This week’s AI story was not just smarter models. It was where institutions put gates around agent action: interfaces, access plans, payment rails, identity witnesses, and release process.
AI labs’ all-or-nothing race leaves no time to fuss about safety
They have ideas about how to restrain wayward models, but worry that doing so will disadvantage them

AI #163: Mythos Quest
There exists an AI model, Claude Mythos, that has discovered critical safety vulnerabilities in every major operating system and browser.

Top AI Security Incidents of 2025 Revealed | Adversa AI
Discover how AI systems are being hacked in the wild — from prompt injection to agent abuse — with real breaches, lessons, and defenses in Adversa AI’s 2025 report.
