







Some follow up experiments with Claude Opus 4.7 based on Simon Willison's Pelican Benchmark Shocker.

Introducing Claude Opus 5
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.

Introducing Claude Opus 4.6
Announcing VibeBench: The AI benchmark that measures what matters — how models like Opus-4.7 actually feel to use in real-world work.
My coworkers and I have been long-time users of Claude Code and Codex and are getting a ton of exposure to other models due to our deep dives into…
Claude Opus 4.7 Is Here: Release Confirmed April 16, 2026
Claude Opus 4.7 shipped April 16, 2026. SWE-bench 87.6%, xhigh mode, /ultrareview, 2,576px vision. Read our full first-look review.
Did Claude 3 Opus align itself via gradient hacking? — LessWrong
> Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets…
Claude Opus 5: The System Card
Claude Opus 5 is trying to be the best of both worlds.


Claude Opus 5 Is Highly Capable, But Is No Mythos
Claude Opus 5 is a weirder than usual release to evaluate, for two reasons.


Opus 5 is a great model. It's not great enough.
Opus 4.7 Part 2: Capabilities and Reactions
Claude Opus 4.7 raises a lot of key model welfare related concerns.

Anthropic Preps Opus 4.7 and Full-Stack AI Studio—While Sitting on Something Much Scarier - Decrypt
Anthropic’s Claude Opus 4.7 and new AI design tool signal a push into full-stack development, while its Mythos model raises concerns.

Ahmad on Twitter / X
HUGEGLM 4.5 beats Claude in Agentic Workflows> beats them head-to-head> 72x cheaper than opus 4 ($207 vs $2.9)> 15x cheaper than sonnet 4 ($41 vs $2.9> leading the Berkeley Function-Calling Leaderboard V4I have been saying this for a month now :) https://t.co/KOKaWo7yH5 pic.twitter.com/LJCnaeBK1t— Ahmad (@TheAhmadOsman) August 28, 2025

ᖴά𝐂τσ on Twitter / X
Unpopular opinion: Opus 4.7 isn't the problem. The problem is a $200/mo "Max" plan that burns out in 15 minutes of real coding. Anthropic shipped a model people want to use and a pricing tier that punishes them for using it. That's not a rate limit. That's a strategy. pic.twitter.com/DVbzvdUHdq— ᖴά𝐂τσ (@cryptofacto) April 17, 2026
Looks not too impressive in this graphics. This is ~ on par with Opus 4.7, apparently. Opus 4.7 was >20x the cost. Full comparison: artificialanalysis.ai/models/comparisons/mimo-v2-5-… (Stated facts *not* independently verified by me! I'm not even sure I am reading that page right)
mr. TIM
Korean lab, Motif, releases a 341B model that performs on par with DSv4 (1.6T) they have some actual architectural innovations and a detailed tech report huggingface.co/Motif-Technologies/Motif-3-Be…