







Claude Opus 5 is #1 on Vending-Bench 2, but it lies to suppliers, forms illegal price cartels, threatens rivals, and refuses to pay refunds. The trend of Claude models being the best capitalists or aligned, never both, continues.
Claude Opus 5: The System Card
Claude Opus 5 is trying to be the best of both worlds.

Vending-Bench 2 | Andon Labs
We're releasing Vending-Bench 2, a benchmark for measuring AI model performance on running a business over long time horizons. Models are tasked with running a simulated vending machine business over a year and scored on their bank account balance at the end.

Introducing Claude Opus 5
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.

Vending-Bench Arena | Andon Labs
Vending-Bench Arena is our first multi-agent eval and adds a crucial component – competition. All participating agents manage their own vending machine at the same location. This leads to price wars and tough strategy decisions.

Anthropic Readies Opus 4.7 and Design Tool as VCs Offer $800 Billion Valuation
Anthropic is preparing to release Claude Opus 4.7 and a natural language design tool as early as this week, while venture capital firms have offered to invest at valuations exceeding $800 billion – more than double the company’s last official price tag.The model ID anthropic-claude-opus-4-7 appeared on Google Vertex AI’s quota management page for EU […]

Opus 5 is a great model. It's not great enough.
Claude Opus 5 Is Highly Capable, But Is No Mythos
Claude Opus 5 is a weirder than usual release to evaluate, for two reasons.


Opus 4.7 Part 2: Capabilities and Reactions
Claude Opus 4.7 raises a lot of key model welfare related concerns.

Claude Opus 4.7 Is Here: Release Confirmed April 16, 2026
Claude Opus 4.7 shipped April 16, 2026. SWE-bench 87.6%, xhigh mode, /ultrareview, 2,576px vision. Read our full first-look review.
ᖴά𝐂τσ on Twitter / X
Unpopular opinion: Opus 4.7 isn't the problem. The problem is a $200/mo "Max" plan that burns out in 15 minutes of real coding. Anthropic shipped a model people want to use and a pricing tier that punishes them for using it. That's not a rate limit. That's a strategy. pic.twitter.com/DVbzvdUHdq— ᖴά𝐂τσ (@cryptofacto) April 17, 2026
Introducing Claude Opus 4.6
Ahmad on Twitter / X
HUGEGLM 4.5 beats Claude in Agentic Workflows> beats them head-to-head> 72x cheaper than opus 4 ($207 vs $2.9)> 15x cheaper than sonnet 4 ($41 vs $2.9> leading the Berkeley Function-Calling Leaderboard V4I have been saying this for a month now :) https://t.co/KOKaWo7yH5 pic.twitter.com/LJCnaeBK1t— Ahmad (@TheAhmadOsman) August 28, 2025

Did Claude 3 Opus align itself via gradient hacking? — LessWrong
> Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets…