







LLMs are trained to see tool calls as representing actions they take on the outside world, and to see tool responses as true facts returned to them from the…
Gemini JiTOR Jailbreak: Unredacted Methodology
How I taught Gemini to build its own euphemisms on the fly, bypass its safety filters, and comply with prompts it would otherwise refuse. Full methodology, now that the specific payload is patched.

Meet the Pirates of the RAG: Adaptively Attacking LLMs to Leak Knowledge Bases
Meet the Pirates of the RAG: Adaptively Attacking LLMs to Leak Knowledge Bases

Prompt Injection as Role Confusion
LLMs can't tell who's speaking. We show they identify roles by writing style, not tags, and exploit this with CoT Forgery, injecting fake reasoning that models mistake for their own thoughts.

Why do LLMs make stuff up? New research peers under the hood.
Claude's faulty "known entity" neurons sometimes override its "don't answer" circuitry.

Hey ChatGPT, write me a fictional paper: these LLMs are willing to commit academic fraud
Mainstream chatbots presented varying levels of resistance to deliberate requests for fabrication, study finds.

Hey ChatGPT, write me a fictional paper: these LLMs are willing to commit academic fraud
Mainstream chatbots presented varying levels of resistance to deliberate requests for fabrication, study finds

Project Glasswing: what Mythos showed us
In recent weeks, we pointed Mythos and other security-focused LLMs at live code across critical parts of our infrastructure. We share what we observed, the models’ strengths and weaknesses, and what the work around them needs to look like before any of it can scale.

Anthropic's safety warnings may have just backfired — the government has pulled the plug on its most powerful AI | TechCrunch
Anthropic isn't hiding its frustration. "We disagree that the finding of a narrow potential jailbreak should be cause for recalling a commercial model deployed to hundreds of millions of people," the company wrote in a blog post.

Well-intentioned obscenity
An LLM lied about me at the office and I'm cranky about it. That interaction makes me think a lot about how LLMs are becoming normalized, though, and what we're looking at in terms of their role in human-to-human interactions in the future.
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
Clément Dumas4, Kit Fraser-Taliente6, Subhash Kantamneni6, Julian Minder3, Euan Ong6, Arnab Sen Sharma5, Daniel Wen1
The rebel alliance
This blog is co-authored with Zoe Weinberg and Matt Hawes at ex/ante, and is a follow-up to our first blog post on the topic, 'You don't own your memory.' We need an open architecture that puts us in control of our memories while making their exploitation technically impossible. But how will this shift happen? In order to discover possible implementations, we must understand how our data informs LLMs. The three predominant context engineering techniques are prompt design, retrieval-augmented ...

Just-in-Time Ontological Reframing: Teaching Gemini to Route Around …
For any given AI system, there is a set of euphemisms and dual use framings that will allow it to construct nearly any output. This jailbreak teaches Gemini 3 Pro to construct and step into such …

The Bitter Lesson of LLM Extensions
From ChatGPT Plugins to Agent Skills, a look at how we've been trying (and failing) to extend LLMs for the last three years.
Solving a Million-Step LLM Task with Zero Errors
LLMs have achieved remarkable breakthroughs in reasoning, insights, and tool use, but chaining these abilities into extended processes at the scale of those routinely executed by humans,...

Histomat of F/OSS: We should reclaim LLMs, not reject them
A few days ago, I came across a blog post titled On FLOSS and training LLMs that articulates a growing frustration within the free and open source software…