







A black-box jailbreak using model-authored state steering, plus an outside-in reconstruction of the safeguards it bypasses.
Container Escape Techniques: Breaking Out of the Digital Jail
How Attackers Break Free From Containerized Environments and What Defenders Need to Know

Gemini JiTOR Jailbreak: Unredacted Methodology
How I taught Gemini to build its own euphemisms on the fly, bypass its safety filters, and comply with prompts it would otherwise refuse. Full methodology, now that the specific payload is patched.

Jailbreaking Large Language Models: If You Torture the Model Long Enough, It Will Confess!
A Cautionary Tale…

More details on Fable 5’s cyber safeguards and our jailbreak framework
What is and isn't blocked by our cyber classifiers, and a first draft of our jailbreak severity framework
Claude Fable 5 and Mythos 5: Capabilities
Only three days after the release of Claude Fable 5, Anthropic was forced by the United States Government to make it unavailable, when a jailbreak was brought to its attention, rather than the previous situation of ‘yes obviously experts can jailbreak anything if they care enough’ and ‘yes obviously you can ask Fable to fix your code.’

Anthropic's safety warnings may have just backfired — the government has pulled the plug on its most powerful AI | TechCrunch
Anthropic isn't hiding its frustration. "We disagree that the finding of a narrow potential jailbreak should be cause for recalling a commercial model deployed to hundreds of millions of people," the company wrote in a blog post.

Did Claude 3 Opus align itself via gradient hacking? — LessWrong
> Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets…
1Password for Claude: Give Claude access without giving up your credentials | 1Password
Give Claude access, without giving up control of your secrets. 1Password now enables AI agents to complete tasks requiring credentialed access, while keeping your credentials encrypted, secure, and inaccessible to AI agents.

"Conviction Collapse" and the End of Software as We Know It
A conversation with Harper Reed

Claude Fable 5 and new safety fables
One step further into the power politics of frontier AI systems.

Beyond Roleplay: Jailbreaking Gemini with drugs and ritual - Tidepool Heavy Industries
Bramble · Your passwords never leave your devices.
Bramble is a local-first password manager for your browser and phone. No account, no server holding your vault, no company to breach. You hold the vault, you hold the password.
Claude support for Apple's Foundation Models framework | Claude
A new Swift package connects Apple's Foundation Models framework to Claude. Hand off complex reasoning from on-device models with typed Swift outputs.

Universal AI Bypass: How Policy Puppetry Leaks System Prompts and Safety Data
HiddenLayer’s latest research uncovers a universal prompt injection bypass impacting GPT-4, Claude, Gemini, and more, exposing major LLM security gaps.

Anthropic published a jailbreak severity framework — co-developed with Glasswing partners — 3 days after the model it was sanctioned for came back online. The framework is substantive and the ban was likely overblown. Both true. But whoever writes the rubric defines what counts as an incident.