







Compare AI model performance on GDPval-AA v2 Leaderboard. GDPval-AA v2 is Artificial Analysis' evaluation framework for OpenAI's GDPval dataset. It tests AI models on real-world tasks across 44 occupations and 9 major industries. Models are given shell access and web browsing capabilities in an agentic loop via Stirrup to solve tasks, with Elo ratings derived from blind pairwise comparisons.
LLM Leaderboard - Comparison of over 100 AI models from OpenAI, Google, DeepSeek & others
Comparison and ranking the performance of over 100 AI models (LLMs) across key metrics including intelligence, price, performance and speed (output speed - tokens per second & latency - TTFT), context window & others.

AI Values Dashboard
How leading AI models value different people, companies, countries, and groups.
WebDev AI Leaderboard - Best AI Models for Web Development
View overall rankings across AI models on front-end web development tasks, including agentic coding workflows that require multi-step reasoning and tool use.

LLM Rankings | OpenRouter
LLM rankings and AI leaderboard based on benchmarks and real usage data from millions of users. See which AI models developers actually use.
AI Model Leaderboards & Benchmarks
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.
AI Index | Stanford HAI
The mission of the AI Index is to provide unbiased, rigorously vetted, and globally sourced data for policymakers, researchers, journalists, executives, and the general public to develop a deeper understanding of the complex field of AI. To achieve this, we track, collate, distill, and visualize dat
AI Coding Agent Benchmarks & Leaderboard | Artificial Analysis
We measure real-world performance of coding agents on software engineering tasks, including cost, token usage, and execution time. We compare how performance changes across agents, models, and execution settings.
Import AI 452: Scaling laws for cyberwar; rising tides of AI automation; and a puzzle over gDP forecasting
How much could AI revolutionize the economy?

AI Economy Institute - Microsoft Research
The AI Economy Institute (AIEI) is Microsoft’s flagship think tank dedicated to shaping an inclusive, trustworthy AI economy. We building a network of scholars and convening that network with our subject matter experts to explore how artificial intelligence is transforming work, education, and productivity – and making this knowledge base available to policy-makers, educators, and […]

LLM Leaderboard - Best Text & Chat AI Models Compared
Compare and explore Text models ranked by overall performance.

AI Leaderboard 2026: Compare & Rank 300+ Top AI Models by Intelligence, Speed & Price
The AI Leaderboard — independent rankings of GPT, Claude, Gemini, Llama, DeepSeek and 300+ AI models by intelligence, speed and price. Composite LLM Stats Score updated continuously from public benchmarks and live API metrics.

Building and evaluating AI Agents — Sayash Kapoor, AI Snake Oil
Find Open Datasets for AI and Research | Kaggle
Browse and download hundreds of thousands of open datasets for AI research, model training, and analysis. Join a community of millions of researchers, developers, and builders to share and collaborate on Kaggle.

The Productivity Is Real. The Scaling Isn't.
What running an AI agent team taught me about why organizations can't do what one person can.

Arena AI: The Official AI Ranking & LLM Leaderboard
Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.
