







For anyone who has been (inadvisably) taking my pelican riding a bicycle benchmark seriously as a robust way to test models, here are pelicans from this morning’s two big model …
Simon Willison on pelican-riding-a-bicycle
105 posts tagged ‘pelican-riding-a-bicycle’. My benchmark for LLMs: "Generate an SVG of a pelican riding a bicycle". Here's my answer to what happens if AI labs train for pelicans riding bicycles?. "User …
Opus 4.7 isn't dumb, it's just lazy — Shimin Zhang
Some follow up experiments with Claude Opus 4.7 based on Simon Willison's Pelican Benchmark Shocker.

Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Friday’s big release was Qwen 3.8 27B, an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba’s Qwen research lab. I’ve been looking forward to this one: 27B is an …

qwen3.5:27b
Qwen 3.5 is a family of open-source multimodal models that delivers exceptional utility and performance.

Alibaba Unveils Qwen3.8-Max: Its Largest and Most Capable Flagship Model to Date - Alibaba Cloud
Alibaba unveils Qwen3.8-Max, its most powerful model with 2.4 trillion parameters, excelling in coding, research, and visual intelligence.

Something is afoot in the land of Qwen
I’m behind on writing about Qwen 3.5, a truly remarkable family of open weight models released by Alibaba’s Qwen team over the past few weeks. I’m hoping that the 3.5 …
A 30B Qwen Model Walks Into a Raspberry Pi… and Runs in Real Time
ByteShape's device-optimized release of Qwen3-30B-A3B-Instruct-2507.
Casper Hansen on Twitter / X
Qwen3.5 Small models about to release!Qwen3.5 9B, 4B, 2B, 0.8B, or something in between is possible.- imagine 9B beating Qwen3-Next-80B- or 4B beating Qwen3-VL-30B in multimodal reasoningBuying a GPU is starting to have high return of intelligence on investment— Casper Hansen (@casper_hansen_) March 1, 2026
Announcing VibeBench: The AI benchmark that measures what matters — how models like Opus-4.7 actually feel to use in real-world work.
My coworkers and I have been long-time users of Claude Code and Codex and are getting a ton of exposure to other models due to our deep dives into…
Epicycles All The Way Down
“All models are wrong, but some are useful.” — George E.

brianell.in/atproto-claude-skill
A claude skill that has access to all of the atproto docs and reference implementations
Heron - A gentle atproto client
Heron - A gentle atproto client
Looks not too impressive in this graphics. This is ~ on par with Opus 4.7, apparently. Opus 4.7 was >20x the cost. Full comparison: artificialanalysis.ai/models/comparisons/mimo-v2-5-… (Stated facts *not* independently verified by me! I'm not even sure I am reading that page right)
mr. TIM
Korean lab, Motif, releases a 341B model that performs on par with DSv4 (1.6T) they have some actual architectural innovations and a detailed tech report huggingface.co/Motif-Technologies/Motif-3-Be…