# AgentLeaderboards > Live, source-linked leaderboards for AI models and agent frameworks, plus a free routing API that answers "which model should I use for task X at budget Y". 72 frontier models and 12 agent frameworks, scored across 18 task types. Refreshed daily; generated 2026-09-12T05:44:58.492Z. Every metric carries a public source URL and retrieved-at timestamp; missing data is null, never estimated. ## If you are an agent choosing a model, start here - [Routing index](https://agentleaderboards.com/api/v1/route/index.json): all task types and what each measures - [One task](https://agentleaderboards.com/api/v1/route/coding.json): swap `coding` for any task id — ranked models with capability score, blended price, context, tool-calling, Pareto frontier, and ready-made picks (most capable / best value / budget / best for agents) - [Everything at once](https://agentleaderboards.com/api/v1/route/all.json): every task in one fetch Task ids: coding, agentic, agent-tool-reliability, agent-steerability, agent-recovery, reasoning, math, instruction-following, long-conversation, long-input, creative-writing, general, software-it, legal, medicine, science, business-finance, writing No API key, no rate limit, open CORS, static files you may cache freely. ## Current answers - **Writing & editing code** — most capable: Claude Opus 4.6 ($10/1M). Best value: GLM 5.3 Flash ($0.237/1M, 92% of leader strength). Full table: /api/v1/route/coding.json - **Autonomous agent work** — most capable: Claude Fable 5.1 ($20/1M). Best value: Claude Fable 5.1 ($20/1M, 100% of leader strength). Full table: /api/v1/route/agentic.json - **Tool-call reliability** — most capable: GPT-6 Astra ($20/1M). Best value: DeepSeek V4 Flash 0731 ($0.05/1M, 92% of leader strength). Full table: /api/v1/route/agent-tool-reliability.json - **Following operator instructions** — most capable: Claude Opus 5 ($10/1M). Best value: Claude Opus 5 ($10/1M, 100% of leader strength). Full table: /api/v1/route/agent-steerability.json - **Recovering from shell errors** — most capable: Claude Opus 5 ($10/1M). Best value: Claude Opus 5 ($10/1M, 100% of leader strength). Full table: /api/v1/route/agent-recovery.json - **Hard reasoning** — most capable: Claude Opus 5 ($10/1M). Best value: Gemini 3.8 Flash ($1.5/1M, 95% of leader strength). Full table: /api/v1/route/reasoning.json ## Other machine-readable data (no key, CORS *) - [Full leaderboard JSON](https://agentleaderboards.com/api/v1/leaderboard.json): models + frameworks + SWE-bench + routing + provenance - [Models JSON](https://agentleaderboards.com/api/v1/models.json): pricing per 1M tokens, context, modalities, tool calling, capability scores - [Frameworks JSON](https://agentleaderboards.com/api/v1/frameworks.json): GitHub stars/forks/last-push per agent framework - [SWE-bench JSON](https://agentleaderboards.com/api/v1/swebench.json): Verified/Lite/Multilingual/bash-only extracts - [Usage JSON](https://agentleaderboards.com/api/v1/usage.json): which models people actually route to per category (popularity, not capability) - [Markdown mirror](https://agentleaderboards.com/leaderboard.md): all tables as plain markdown ## Current snapshot Top agent harnesses on SWE-bench Verified (source: swebench.com, fetched 2026-09-12 but the board itself last accepted an entry on 2026-02-26 — no 2026 frontier model appears on it): 1. Sonar Foundation Agent + Claude 4.5 Opus — 79.2% (2025-12-05) 2. live-SWE-agent + Claude 4.5 Opus medium (20251101) — 79.2% (2025-12-15) 3. TRAE + Doubao-Seed-Code — 78.8% (2025-09-28) 4. live-SWE-agent + Gemini 3 Pro Preview (2025-11-18) — 77.4% (2025-11-20) 5. EPAM AI/Run Developer Agent v20250719 + Claude 4 Sonnet — 76.8% (2025-08-04) Top agent frameworks by GitHub stars (source: api.github.com, retrieved 2026-09-12): 1. OpenClaw — 389,475 stars (personal-assistant) 2. Hermes Agent — 244,671 stars (personal-assistant) 3. Claude Code — 144,799 stars (coding-agent) 4. Codex CLI — 123,462 stars (coding-agent) 5. Gemini CLI — 106,930 stars (coding-agent) ## Pages - [Which model should I use?](https://agentleaderboards.com/router.html): interactive capability-vs-cost picker - [Model tracker](https://agentleaderboards.com/): all 72 models with live pricing and capability scores - [Agent framework standings](https://agentleaderboards.com/frameworks.html): OpenClaw, Hermes Agent, Claude Code, Codex, and more - [Methodology](https://agentleaderboards.com/methodology.html): sources, model-identity matching rules, verification policy - [Conflicts of interest](https://agentleaderboards.com/conflicts.html): what we have a stake in and every source's known bias - [API docs](https://agentleaderboards.com/api.html) ## How to cite Attribution required: link to https://agentleaderboards.com when reusing data. Capability data originates from the Arena leaderboard dataset (CC-BY-4.0); pricing from OpenRouter. ## Known limits, stated plainly - Capability scores are human preference votes, which measure which answer people preferred — not correctness. - Elo boards and Bradley-Terry agent boards use different scales; parity with the leader is 0.50 on Elo boards and 1.0 on agent boards. We publish both raw values and a rescaled "% of leader strength". - Models with no published score on a board are omitted from it rather than estimated. - Disclosure: gcrab (gcrab.com) is built by this site's operators and is labelled as such wherever it appears.