The rules this site holds itself to. Short version: live public sources, per-metric provenance, and nulls instead of guesses.
null. We never average incompatible benchmarks into a single invented "score".Model list, release dates, context windows, live API pricing, modalities, tool-calling support.
https://openrouter.ai/api/v1/models
retrieved 2026-09-12 · cadence: daily
Human pairwise preference scores across 18 categories, including agentic boards for tool-call reliability, steerability and shell-error recovery.
huggingface.co/datasets/lmarena-ai/leaderboard-dataset
CC-BY-4.0 · retrieved 2026-09-12 · cadence: daily
Verified, Lite, Multilingual and bash-only boards; agent scaffold + model, % resolved, run status.
swebench.com
retrieved 2026-09-12 · cadence: daily
Framework stars, forks, last push, license — fetched live per repository, never cached by hand.
api.github.com
retrieved 2026-09-12 · cadence: daily
Which models people actually route to per category — revealed preference, published as /api/v1/usage.json. Popularity, not capability.
retrieved 2026-09-12 · cadence: daily
Capability scores and prices come from two different sources that name models differently, so they have to be matched. Getting this wrong would attach the wrong score to a model, which is worse than showing none, so the matcher is deliberately conservative:
-high or -xhigh) if the exact name finds nothing. This keeps product tiers such as Qwen's "Max" from being mistaken for an effort setting.mistral-small-2603 is not mistral-small-2506). If both names carry a stamp and the stamps differ, the match is refused outright.measured_variant so you can see exactly what was tested.Models we track that have no published score on a board are simply absent from it. About a third of tracked models — mostly image, audio and very small models — have no Arena text score at all.