







Quantitative benchmarks of LLM code editing skill.
Best LLM for Coding 2026 | AI Coding Model Rankings & Benchmarks
Which AI model writes the best code? We rank every major LLM — open and closed source — across SWE-bench, HumanEval, LiveCodeBench, and Terminal-Bench coding benchmarks. Compare the best LLMs for coding, software engineering, and programming.

LLM Leaderboard 2026 — Compare Top AI Models
Compare the latest LLM benchmarks for GPT, Claude, Gemini and more. Updated rankings across reasoning, coding, math, and multilingual tasks with pricing and speed data.
Wolfram LLM Benchmarking Project
Results from Wolfram's ongoing tracking of LLM performance. The benchmark is based on a Wolfram Language code generation task.

AI Model Leaderboards & Benchmarks
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.

LLM Leaderboard - Comparison of over 100 AI models from OpenAI, Google, DeepSeek & others
Comparison and ranking the performance of over 100 AI models (LLMs) across key metrics including intelligence, price, performance and speed (output speed - tokens per second & latency - TTFT), context window & others.

LLM Rankings | OpenRouter
LLM rankings and AI leaderboard based on benchmarks and real usage data from millions of users. See which AI models developers actually use.
Berkeley Function Calling Leaderboard (BFCL) V4
Explore The Berkeley Function Calling Leaderboard (also called The Berkeley Tool Calling Leaderboard) to see the LLM's ability to call functions (aka tools) accurately.
AI Leaderboard 2026: Compare & Rank 300+ Top AI Models by Intelligence, Speed & Price
The AI Leaderboard — independent rankings of GPT, Claude, Gemini, Llama, DeepSeek and 300+ AI models by intelligence, speed and price. Composite LLM Stats Score updated continuously from public benchmarks and live API metrics.

The Kaitchup Index: A Leaderboard for LLMs and Their Quantized Versions
Comparing formats like GGUF, GPTQ, and AWQ, with different bitwidths

Honey, I Shrunk the Coding Agent
Coding Agent Adaptation Lets a 9B LLM Outperform 10x Larger Models on Aider Polyglot Benchmark

AI Coding Agent Benchmarks & Leaderboard | Artificial Analysis
We measure real-world performance of coding agents on software engineering tasks, including cost, token usage, and execution time. We compare how performance changes across agents, models, and execution settings.
The Case Against LLMs as Rerankers
Authors: Apoorva Joshi, Zhenmei Shi, Akshay Goindani, Hong LiuResearch Leads: Zhenmei Shi, Akshay Goindani, Hong Liu Large language models are increasingly being used for a broad range of tasks, in…

ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Recent large language models (LLMs) advancements sparked a growing research interest in tool assisted LLMs solving real-world challenges…
