







The Epoch Capabilities Index combines many benchmarks into a single capability scale for comparing models over time.
Capability Trees: A Protocol-Level Extension of Object Capabilities, Draft 2
This draft addresses questions and feedback from the ATProto community, in particular from Zicklag (@zicklag.dev) and Brooklyn Zelenka (@expede.wtf). Draft 1 remains available for transparency.
Performance Explorer — oMLX
Compare model performance across context lengths with community benchmark data.

Model Performance Evaluation to SubQ 1.1 Small Preview Performance Evaluation | Appen
A third-party benchmark assessment of Subquadratic's preview models, conducted by Appen across long-context retrieval, code generation, business-workflow automation, and graduate-level reasoning benchmarks.
Capability Trees: A Protocol-Level Extension of Object Capabilities, Draft 3
This draft incorporates feedback from Brooklyn Zelenka (@expede.wtf), Daniel Holmgren (@dholms.at), and Zicklag (@zicklag.dev), and addresses questions raised in the ATProto Private Data Working Group. Draft 2 remains available for transparency.
Capability Trees: A Protocol-Level Extension of Object Capabilities, Draft 4
This draft incorporates feedback from Brooklyn Zelenka (@expede.wtf) and Daniel Holmgren (@dholms.at), and developments in the ATProto Private Data Working Group. Previous drafts remain available for transparency.
Object-capability model
The object-capability model is a computer security model. A capability describes a transferable right to perform one (or more) operations on a given object. It can be obtained by the following combination:
What is a Capability?
I have several folks asking me what I mean when I say capability in the context of what we are building with Naftiko. I have been writing about API capabilit...

LLM Leaderboard - Comparison of over 100 AI models from OpenAI, Google, DeepSeek & others
Comparison and ranking the performance of over 100 AI models (LLMs) across key metrics including intelligence, price, performance and speed (output speed - tokens per second & latency - TTFT), context window & others.

AI Model & API Providers Analysis | Artificial Analysis
Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.

GDPval-AA v2 Leaderboard | Artificial Analysis
Compare AI model performance on GDPval-AA v2 Leaderboard. GDPval-AA v2 is Artificial Analysis' evaluation framework for OpenAI's GDPval dataset. It tests AI models on real-world tasks across 44 occupations and 9 major industries. Models are given shell access and web browsing capabilities in an agentic loop via Stirrup to solve tasks, with Elo ratings derived from blind pairwise comparisons.
The Kaitchup Index: A Leaderboard for LLMs and Their Quantized Versions
Comparing formats like GGUF, GPTQ, and AWQ, with different bitwidths

ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Recent large language models (LLMs) advancements sparked a growing research interest in tool assisted LLMs solving real-world challenges…

Open-world evaluations for measuring frontier AI capabilities
Introducing CRUX, a new project for evaluating AI on long, messy tasks

MirrorCode: A benchmark for real-world software projects