







You need a lot of data points to understand a new model, and what you have.
Opus 4.7 Part 2: Capabilities and Reactions
Claude Opus 4.7 raises a lot of key model welfare related concerns.

Announcing VibeBench: The AI benchmark that measures what matters — how models like Opus-4.7 actually feel to use in real-world work.
My coworkers and I have been long-time users of Claude Code and Codex and are getting a ton of exposure to other models due to our deep dives into…
Introducing Claude 4
Discover Claude 4's breakthrough AI capabilities. Experience more reliable, interpretable assistance for complex tasks across work and learning.
Opus 4.7 isn't dumb, it's just lazy — Shimin Zhang
Some follow up experiments with Claude Opus 4.7 based on Simon Willison's Pelican Benchmark Shocker.

Introducing Claude Opus 4.6
Better Models: Worse Tools
About an aggravating tool-calling regression in newer Claude models.

Models overview
Claude is a family of state-of-the-art large language models developed by Anthropic. This guide introduces the available models and compares their performance.
Opus 4.7 Part 3: Model Welfare
It is thanks to Anthropic that we get to have this discussion in the first place.

Anthropic Preps Opus 4.7 and Full-Stack AI Studio—While Sitting on Something Much Scarier - Decrypt
Anthropic’s Claude Opus 4.7 and new AI design tool signal a push into full-stack development, while its Mythos model raises concerns.

Introduction - How to Write an Inference Engine
A zero-to-hero guide to Muse Glimmer on Apple Metal, kvpack, and disaggregated NVFP4 prefill.

I Think They Are Lying To You
Anthropic Readies Opus 4.7 and Design Tool as VCs Offer $800 Billion Valuation
Anthropic is preparing to release Claude Opus 4.7 and a natural language design tool as early as this week, while venture capital firms have offered to invest at valuations exceeding $800 billion – more than double the company’s last official price tag.The model ID anthropic-claude-opus-4-7 appeared on Google Vertex AI’s quota management page for EU […]

Introducing Claude Opus 5
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.

Opus 5 is a great model. It's not great enough.
Looks not too impressive in this graphics. This is ~ on par with Opus 4.7, apparently. Opus 4.7 was >20x the cost. Full comparison: artificialanalysis.ai/models/comparisons/mimo-v2-5-… (Stated facts *not* independently verified by me! I'm not even sure I am reading that page right)
mr. TIM
Korean lab, Motif, releases a 341B model that performs on par with DSv4 (1.6T) they have some actual architectural innovations and a detailed tech report huggingface.co/Motif-Technologies/Motif-3-Be…