







ᖴά𝐂τσ on Twitter / X
Unpopular opinion: Opus 4.7 isn't the problem. The problem is a $200/mo "Max" plan that burns out in 15 minutes of real coding. Anthropic shipped a model people want to use and a pricing tier that punishes them for using it. That's not a rate limit. That's a strategy. pic.twitter.com/DVbzvdUHdq— ᖴά𝐂τσ (@cryptofacto) April 17, 2026
Opus 4.5 Limits have been cracked down on
30 votes, 63 comments. I'm on the Pro plan, but sent my last message at midnight last night. Came on at 8:30 am and sent 2 messages, which resulted…

Introducing Claude Opus 5
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.

Local Qwen isn't a worse Opus, it's a different tool
We've all heard people say that Qwen is near-Opus level, but I have receipts and am here to be transparent with you.
Opus 4.7 isn't dumb, it's just lazy — Shimin Zhang
Some follow up experiments with Claude Opus 4.7 based on Simon Willison's Pelican Benchmark Shocker.

How To Argue With An AI Booster
Editor's Note: For those of you reading via email, I recommend opening this in a browser so you can use the Table of Contents. This is my longest newsletter - a 16,000-word-long opus - and if you like it, please subscribe to my premium newsletter. Thanks for reading! In

Opus 5 is a great model. It's not great enough.
Announcing VibeBench: The AI benchmark that measures what matters — how models like Opus-4.7 actually feel to use in real-world work.
My coworkers and I have been long-time users of Claude Code and Codex and are getting a ton of exposure to other models due to our deep dives into…
Did Claude 3 Opus align itself via gradient hacking? — LessWrong
> Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets…
Context Rot: How Increasing Input Tokens Impacts LLM Performance
Large Language Models (LLMs) are typically presumed to process context uniformly—that is, the model should handle the 10,000th token just as reliably as the 100th. However, in practice, this assumption does not hold. We observe that model performance varies significantly as input length changes, even on simple tasks. In this report, we evaluate 18 LLMs, including the state-of-the-art GPT-4.1, Claude 4, Gemini 2.5, and Qwen3 models. Our results reveal that models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows.

Introducing @@ — @@ for Mac
Your agent was never the bottleneck. The trip to it was. Type @@ anywhere on your Mac, and the answer comes back where you were typing.

Ahmad on Twitter / X
HUGEGLM 4.5 beats Claude in Agentic Workflows> beats them head-to-head> 72x cheaper than opus 4 ($207 vs $2.9)> 15x cheaper than sonnet 4 ($41 vs $2.9> leading the Berkeley Function-Calling Leaderboard V4I have been saying this for a month now :) https://t.co/KOKaWo7yH5 pic.twitter.com/LJCnaeBK1t— Ahmad (@TheAhmadOsman) August 28, 2025

kwindla on Twitter / X
Local voice AI with a 235 billion parameter LLM. ✅- smart-turn v2- MLX Whisper (large-v3-turbo-q4)- Qwen3-235B-A22B-Instruct-2507-3bit-DWQ- KokoroAll models running local on an M4 mac. Max RAM usage ~110GB.Voice-to-voice latency is ~950ms. There are a couple of… pic.twitter.com/iYNQlb9JkI— kwindla (@kwindla) July 27, 2025
Looks not too impressive in this graphics. This is ~ on par with Opus 4.7, apparently. Opus 4.7 was >20x the cost. Full comparison: artificialanalysis.ai/models/comparisons/mimo-v2-5-… (Stated facts *not* independently verified by me! I'm not even sure I am reading that page right)
mr. TIM
Korean lab, Motif, releases a 341B model that performs on par with DSv4 (1.6T) they have some actual architectural innovations and a detailed tech report huggingface.co/Motif-Technologies/Motif-3-Be…