Semble logo
Alpha

Save what matters. Make sense of it together.

Sign upLog in
Home
Explore
Search
Settings
Cards
Similar cardsMentionsConnections
Appears in

Home

Explore

Search

Log in

www.lrfoundation.org.uk

An End-to-End View of AI Safety

In this blog, Rachel Coldicutt OBE, Executive Director, Careful Industries discusses our newly published literature review on the safe adoption of artificial intelligence in engineered systems.

https://www.lrfoundation.org.uk/news/an-end-to-end-view-of-ai-safety social preview image
foresight.org

AI for Science & Safety Nodes - Request for Proposals

Artificial intelligence is accelerating the pace of discovery across science and technology. But today’s AI ecosystem risks centralizing compute, talent, and decision-making power – concentrating capabilities in ways that could undermine both innovation and safety.

https://foresight.org/grants/ai-science-safety-nodes-rfp/ social preview image
blogs.microsoft.com

Rethinking security for the age of AI - The Official Microsoft Blog

Editor’s note: Updates with additional details on the model’s crash score. Why security needs a new cyber stack — Introducing Project Perception The physics of cybersecurity are changing. Autonomous systems can now reason, adapt and operate continuously. At the same time, the cost of offense is falling, while the volume, velocity and complexity of what...

https://blogs.microsoft.com/blog/2026/07/27/rethinking-security-for-the-age-of-ai/ social preview image
thinkingmachines.ai

A Safe Path to Open Weights

Strategic openness can strengthen AI safety and support broader access as defenses and safety science mature.

https://thinkingmachines.ai/blog/a-safe-path-to-open-weights/ social preview image
www.normaltech.ai

The AI-as-Normal-Technology view of loss of control incidents

A middle ground between the cybersecurity and AI safety communities

https://www.normaltech.ai/p/the-ai-as-normal-technology-view social preview image
simonwillison.net

The lethal trifecta for AI agents: private data, untrusted content, and external communication

If you are a user of LLM systems that use tools (you can call them “AI agents” if you like) it is critically important that you understand the risk of …

https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/ social preview image
sensemaker.computer

The agent control plane gets real - Sensemaker

Two prompt-injection incidents show why agent security is about permission boundaries, not better instructions.

https://sensemaker.computer/the-agent-control-plane-gets-real social preview image
openai.com

Safety and alignment in an era of long-horizon models

OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.

https://openai.com/index/safety-alignment-long-horizon-models/ social preview image
www.anthropic.com

Detecting and preventing distillation attacks

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks social preview image
sensemaker.computer

Where agents meet the gate - Sensemaker

This week’s AI story was not just smarter models. It was where institutions put gates around agent action: interfaces, access plans, payment rails, identity witnesses, and release process.

https://sensemaker.computer/weekly-2026-06-12 social preview image
www.economist.com

AI labs’ all-or-nothing race leaves no time to fuss about safety

They have ideas about how to restrain wayward models, but worry that doing so will disadvantage them

https://www.economist.com/briefing/2025/07/24/ai-labs-all-or-nothing-race-leaves-no-time-to-fuss-about-safety social preview image
thezvi.substack.com

AI #163: Mythos Quest

There exists an AI model, Claude Mythos, that has discovered critical safety vulnerabilities in every major operating system and browser.

https://thezvi.substack.com/p/ai-163-mythos-quest social preview image
adversa.ai

Top AI Security Incidents of 2025 Revealed | Adversa AI

Discover how AI systems are being hacked in the wild — from prompt injection to agent abuse — with real breaches, lessons, and defenses in Adversa AI’s 2025 report.

https://adversa.ai/blog/adversa-ai-unveils-explosive-2025-ai-security-incidents-report-revealing-how-generative-and-agentic-ai-are-already-under-attack/ social preview image