







AI models have likely reached parity with superforecasters on ForecastBench
New results from the ForecastBench leaderboard suggest the gap between frontier AI and elite human forecasters is closing.

Public LLM exceeds superforecaster on Forecast bench by EOY 2026?
41% chance. Resolves to YES if any LLM released in 2026 exceeds the superforecaster baseline on ForecastBench by July 2027. Resolves to NO if this does not happen, or if after January 1, 2027 we have results from enough LLMs (e.g. the leading models from the major AI labs at the time) to be confident this will not occur. If The Forecasting Research Institute tells us how this market should resolve, then we will go with what they say.

Georgi Gerganov on Twitter / X
llama.cpp at 100k starsnow that 90% of the code worldwide is being written by AI agents, I predict that within 3-6 months, 90% of all AI agents will be running locally with llama.cpp 😄Jokes aside, I am going to use this small milestone as an opportunity to reflect a bit on… pic.twitter.com/BpN5VL1bYV— Georgi Gerganov (@ggerganov) March 30, 2026
Zhijing Jin on Twitter / X
🚀 Can #AI agents actually do science?Current agents can optimize for predictive fit — but fail at recovering the underlying physical laws.We introduce Stargazer🪐: a scalable benchmark + environment for astronomical discovery🔭Agents must:• propose hypotheses•… pic.twitter.com/hGIjz4KJMT— Zhijing Jin (@ZhijingJin) May 7, 2026

KillBench: Discovering Hidden Biases of LLMs
1M+ experiments exposing bias in critical AI decision-making

State of AI 2025: 100T Token LLM Usage Study | OpenRouter
Read OpenRouter's 2025 State of AI report — an empirical 100 trillion token study of real LLM usage, model trends, and developer insights.
Suhail on Twitter / X
Oh, this is coming back in such a big way this year with AI. We were a bit too early.Remember: Browser = OS pic.twitter.com/YoMaS8cjmI— Suhail (@Suhail) June 23, 2025


What the hell happened with AGI timelines in 2026?
Lexicon: How China talks about 'agentic AI' - DigiChina
Three months after the Chinese AI company DeepSeek shocked global markets with a highly capable reasoning model, another China-linked company made a splash with a capable agentic AI system. Did Manus, released in March 2025, portend Chinese leadership in AI systems that go beyond chatbots to take action on the user’s behalf? Victor Mustar, head […]

LLM Leaderboard 2026 — Compare Top AI Models
Compare the latest LLM benchmarks for GPT, Claude, Gemini and more. Updated rankings across reasoning, coding, math, and multilingual tasks with pricing and speed data.
LLM Leaderboard - Comparison of over 100 AI models from OpenAI, Google, DeepSeek & others
Comparison and ranking the performance of over 100 AI models (LLMs) across key metrics including intelligence, price, performance and speed (output speed - tokens per second & latency - TTFT), context window & others.

Import AI 446: Nuclear LLMs; China's big AI benchmark; measurement and AI policy
Will AIs be jealous of one another?

AI Leaderboard 2026: Compare & Rank 300+ Top AI Models by Intelligence, Speed & Price
The AI Leaderboard — independent rankings of GPT, Claude, Gemini, Llama, DeepSeek and 300+ AI models by intelligence, speed and price. Composite LLM Stats Score updated continuously from public benchmarks and live API metrics.
