







Anthropic's safety warnings may have just backfired — the government has pulled the plug on its most powerful AI | TechCrunch
Anthropic isn't hiding its frustration. "We disagree that the finding of a narrow potential jailbreak should be cause for recalling a commercial model deployed to hundreds of millions of people," the company wrote in a blog post.

More details on Fable 5’s cyber safeguards and our jailbreak framework
What is and isn't blocked by our cyber classifiers, and a first draft of our jailbreak severity framework
Why Anthropic’s new model has cybersecurity experts rattled
The company says it has built its most dangerous model yet. Can its coalition of internet companies fix the internet before others catch up?

Developer Ecosystems for Software Safety
How to design and implement information systems so they are safe and secure is a complex topic. Both high-level design principles and implementation guidance for software safety and security are well established and broadly accepted. For example, Jerome Saltzer and Michael Schroeder’s seminal overview of principles of secure design was published almost 50 years ago,10 and various community and governmental bodies have published comprehensive best practices about how to avoid common software weaknesses—for example, Common Weakness Enumeration (CWE)a and Open Worldwide Application Security Project (OWASP) Cheat Sheet Series.b

Why Anthropic believes its latest model is too dangerous to release
“The language models we have now are probably the most significant thing to happen in security since we got the Internet.”

Updates to Consumer Terms and Privacy Policy
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Rethinking security for the age of AI - The Official Microsoft Blog
Editor’s note: Updates with additional details on the model’s crash score. Why security needs a new cyber stack — Introducing Project Perception The physics of cybersecurity are changing. Autonomous systems can now reason, adapt and operate continuously. At the same time, the cost of offense is falling, while the volume, velocity and complexity of what...

Anthropic Drops Flagship Safety Pledge
In an abrupt shift, the company may release future AI models without ironclad safety guarantees

Safety overview: GPT-6 Astra
GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.

Apple's Assault on Standards - Infrequently Noted
By subverting the voluntary nature of open standards, Apple has defanged them as tools that users can employ against the totalising power of native apps in their digital lives. This high-modernist approach is antithetical to the foundational commitments of internet standards bodies and, over time, erode them.
Open by Design: ROOST's Approach to Safety Tool Development

Building the safety foundations for India’s agentic future
Last year, through Google’s Safety Charter for India’s AI-led transformation, we set out our commitment to building trust into India’s digital growth. We also showed how…

A threat model for accessibility on the web - Alice
A explanation of the primary threat to accessibility on the web, and a call to action for the web standards community

Anthropic published a jailbreak severity framework — co-developed with Glasswing partners — 3 days after the model it was sanctioned for came back online. The framework is substantive and the ban was likely overblown. Both true. But whoever writes the rubric defines what counts as an incident.
Why does every platform rebuild the same safety tools from scratch, behind closed doors? Our Head of Product @julietshen.bsky.social joined the Won't Fix pod to talk open-source T&S infrastructure, AI vs. human judgment & safety for the decentralized web: youtube.com/watch?v=RxFV7VwxkLs
Won't Fix Episode 9: With Juliet Shen, Cofounder & HOP at ROOST
www.youtube.com