







I ran this in progressive mode w/Opus4.6 on a paper where refine.ink had given me very useful feedback. It caught many of the same minor issues! Main difference was in high level feedback-with refine it was much closer to what I'd expect from experts. But not bad for a fraction of refine cost! (~$3)
chenhaotan.bsky.social
Peer review is facing a death spiral, and AI production tools are speeding it up. AI-assisted reviewing is necessary and should be open. We built OpenAIReview: open AI reviewing for everyone, for the cost of a coffee. openaireview.github.io/blog.html 🧵
Mar 9, 2026 at 8:08 PM
Opus 5 is a great model. It's not great enough.
Opus 4.7 isn't dumb, it's just lazy — Shimin Zhang
Some follow up experiments with Claude Opus 4.7 based on Simon Willison's Pelican Benchmark Shocker.

Introducing Claude Opus 5
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.

Announcing VibeBench: The AI benchmark that measures what matters — how models like Opus-4.7 actually feel to use in real-world work.
My coworkers and I have been long-time users of Claude Code and Codex and are getting a ton of exposure to other models due to our deep dives into…
Local Qwen isn't a worse Opus, it's a different tool
We've all heard people say that Qwen is near-Opus level, but I have receipts and am here to be transparent with you.
ᖴά𝐂τσ on Twitter / X
Unpopular opinion: Opus 4.7 isn't the problem. The problem is a $200/mo "Max" plan that burns out in 15 minutes of real coding. Anthropic shipped a model people want to use and a pricing tier that punishes them for using it. That's not a rate limit. That's a strategy. pic.twitter.com/DVbzvdUHdq— ᖴά𝐂τσ (@cryptofacto) April 17, 2026
Refine — AI Verification Trusted by World-Class Experts
Good decisions require verified quality. Refine devotes hours of frontier compute to protect your work and reputation from fixable mistakes.

Aniket Panjwani on Twitter / X
Here's something I don't understand about Refine.1 review costs $50 and 10 reviews cost $300. So, unless they're taking a loss, average cost of inference on a Refine review is no more than $30.On a ChatGPT Pro plan, for $200/mo, you're getting the equivalent of ~$14k/mo in…— Aniket Panjwani (@aniketapanjwani) August 20, 2026
Did Claude 3 Opus align itself via gradient hacking? — LessWrong
> Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets…
Prompting Claude Opus 5
Behavioral differences and prompting patterns for Claude Opus 5, covering response verbosity, agentic narration, task scoping, subagent delegation, self-correction, and output artifacts when thinking is disabled.

Refine
I recently tried refine, an AI tool for refining academic articles, developed by Yann Calvó López and Ben Golub.

Claude Opus 5: The System Card
Claude Opus 5 is trying to be the best of both worlds.

Opus 4.5 Limits have been cracked down on
30 votes, 63 comments. I'm on the Pro plan, but sent my last message at midnight last night. Came on at 8:30 am and sent 2 messages, which resulted…
How To Argue With An AI Booster
Editor's Note: For those of you reading via email, I recommend opening this in a browser so you can use the Table of Contents. This is my longest newsletter - a 16,000-word-long opus - and if you like it, please subscribe to my premium newsletter. Thanks for reading! In

Looks not too impressive in this graphics. This is ~ on par with Opus 4.7, apparently. Opus 4.7 was >20x the cost. Full comparison: artificialanalysis.ai/models/comparisons/mimo-v2-5-… (Stated facts *not* independently verified by me! I'm not even sure I am reading that page right)
mr. TIM
Korean lab, Motif, releases a 341B model that performs on par with DSv4 (1.6T) they have some actual architectural innovations and a detailed tech report huggingface.co/Motif-Technologies/Motif-3-Be…