







Announcing VibeBench: The AI benchmark that measures what matters — how models like Opus-4.7 actually feel to use in real-world work.
My coworkers and I have been long-time users of Claude Code and Codex and are getting a ton of exposure to other models due to our deep dives into…
GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index
Benchmarks and Analysis of GLM-5.2

zai-org/GLM-5.3-Flash · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
unsloth/GLM-5.2-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
We introduce SimpleQA Verified, a 1,000-prompt benchmark for evaluating Large Language Model (LLM) short-form factuality based on OpenAI's SimpleQA. It addresses critical limitations in OpenAI's benchmark, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created through a rigorous multi-stage filtering process involving de-duplication, topic balancing, and source reconciliation to produce a more reliable and challenging evaluation set, alongside improvements in the autorater prompt. On this new benchmark, Gemini 2.5 Pro achieves a state-of-the-art F1-score of 55.6, outperforming other frontier models, including GPT-5. This work provides the research community with a higher-fidelity tool to track genuine progress in parametric model factuality and to mitigate hallucinations. The benchmark dataset, evaluation code, and leaderboard are available at: https://www.kaggle.com/benchmarks/deepmind/simpleqa-verified.

SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
We introduce SimpleQA Verified, a 1,000-prompt benchmark for evaluating Large Language Model (LLM) short-form factuality based on OpenAI's SimpleQA. It addresses critical limitations in OpenAI's benchmark, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created through a rigorous multi-stage filtering process involving de-duplication, topic balancing, and source reconciliation to produce a more reliable and challenging evaluation set, alongside improvements in the autorater prompt. On this new benchmark, Gemini 2.5 Pro achieves a state-of-the-art F1-score of 55.6, outperforming other frontier models, including GPT-5. This work provides the research community with a higher-fidelity tool to track genuine progress in parametric model factuality and to mitigate hallucinations. The benchmark dataset, evaluation code, and leaderboard are available at: https://www.kaggle.com/benchmarks/deepmind/simpleqa-verified.

Open Models Inference for Coding · Umans AI
Hosted Kimi K3, GLM 5.2, and DeepSeek V4 Flash. Pay per token, on infrastructure we own.


Opus 5 is a great model. It's not great enough.
GLM Coding Plan — AI Coding Powered by GLM-5.1 & GLM-5-Turbo for Agents & IDEs
Use GLM models like GLM-5.1 & GLM-5-Turbo for AI coding in Claude Code, Kilo Code, Cline, OpenCode, Clawdbot/OpenClaw and more. Plans from 18/month—fast, reliable code generation and tool use for daily dev work.
Prime Agent: A self-improving RLM agent
Prime Agent is our open-source, self-improving coding harness built around two abstractions: the Recursive Language Model (RLM) and the Continual Harness. With Opus 5, it achieves 95.5% on ARC-AGI-3, surpassing the reported human expert baseline.

GLM-OCR Technical Report
GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-parameter GLM...

https://z.ai/blog/glm-4.5
mradermacher/GLM-4.6V-Flash-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.

Simon Willison on Twitter / X
I think it's non-obvious to many people that the OpenAI voice mode runs on a much older, much weaker model - it feels like the AI that you can talk to should be the smartest AI but it really isn't https://t.co/bZ0Qqx9Sa9— Simon Willison (@simonw) April 10, 2026
GLM-5.2 - open weights - single-shot our AI-resistant backend take-home to a higher level than Opus 4.8, and built offmute-v2: state-of-the-art timestamp-accurate diarization. A head-to-head with no detail glossed over.