Which AI model or agent should you use?

Live standings for 72 frontier models and 12 agent frameworks β€” capability measured on the kind of work you're actually doing, priced against what it costs per million tokens. Every number links to its public source, and anything we can't source is left blank instead of estimated.

Auto-refreshed daily Β· last update 2026-09-12 05:44 UTC

Straight answers

The picks below come from measured human-preference scores on each kind of work, divided against live API pricing. Change the task or set a budget β†’

Best model for coding

Writing and editing code
Claude Opus 4.6
Anthropic Β· $10.00/1M blended
Value pick: GLM 5.3 Flash at $0.24/1M β€” 98% cheaper, 92% of leader strength.

Best model for agents

Autonomous multi-step tool use
Claude Fable 5.1
Anthropic Β· $20.00/1M blended
Nothing cheaper stays within 90% of it on this board.

Best all-round model

General assistant work
Claude Fable 5.1
Anthropic Β· $20.00/1M blended
Value pick: Gemini 3.8 Flash at $1.50/1M β€” 93% cheaper, 95% of leader strength.

New releases last 3 weeks

ModelProviderReleasedContextInput /1MOutput /1M
DeepSeek V4.1 FlashDeepSeek2026-09-10 new1.048576M$0.15$0.60
GPT-6 AstraOpenAI2026-09-04 new1.05M$10.00$50.00
GPT-6 Astra ProOpenAI2026-09-04 new1.05M$10.00$50.00
Qwen3.8 Max (0902)Qwen (Alibaba)2026-09-03 new1M$2.00$6.00
Gemini 3.8 FlashGoogle2026-09-02 new1.048576M$0.75$3.75
Claude Fable 5.1Anthropic2026-09-01 new1M$10.00$50.00
Qwen3.8 FlashQwen (Alibaba)2026-08-26 new1M$0.15$0.47
GLM 5.3 FlashZ.ai (Zhipu)2026-08-26 new1.31072M$0.15$0.50

Source: OpenRouter model catalog Β· retrieved 2026-09-12

All tracked models 72

General and Coding are Arena Elo for that category β€” human preference votes, higher is better. AA is the Artificial Analysis Intelligence Index, a benchmark-based composite on a different scale; it's shown beside the Elo columns as an independent cross-check, never blended with them. Blank means no published score, never an estimate. Price is blended (inputΓ—3 + outputΓ—1) Γ· 4 per million tokens; 18 models tagged varies are priced differently by UTC day or hour, and the full schedule is in the API. Click any header to sort.

ModelProviderGeneralCodingAA$/1MContextToolsReleased
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash
DeepSeek β€” β€” 39.5 $0.26 varies 1.048576M tools 2026-09-10
GPT-6 Astra
openai/gpt-6-astra
OpenAI 1442 1485 52.8 $20.00 varies 1.05M tools 2026-09-04
GPT-6 Astra Pro
openai/gpt-6-astra-pro
OpenAI β€” β€” β€” $20.00 varies 1.05M tools 2026-09-04
Qwen3.8 Max (0902)
qwen/qwen3.8-max-0902
Qwen (Alibaba) 1481 1501 40.3 $3.00 1M tools 2026-09-03
Gemini 3.8 Flash
google/gemini-3.8-flash
Google 1494 1514 41.2 $1.50 1.048576M tools 2026-09-02
Claude Fable 5.1
anthropic/claude-fable-5.1
Anthropic 1510 1512 53.4 $20.00 1M tools 2026-09-01
Qwen3.8 Flash
qwen/qwen3.8-flash
Qwen (Alibaba) β€” β€” β€” $0.23 1M tools 2026-08-26
GLM 5.3 Flash
z-ai/glm-5.3-flash
Z.ai (Zhipu) 1472 1507 41.9 $0.24 1.31072M tools 2026-08-26
DeepSeek V4 Flash Vision Exp
deepseek/deepseek-v4-flash-vision-exp
DeepSeek β€” β€” β€” $0.33 1.048576M tools 2026-08-21
GLM 5.3
z-ai/glm-5.3
Z.ai (Zhipu) 1477 1502 44.9 $2.15 1.31072M tools 2026-08-18
Qwen3.8 27B
qwen/qwen3.8-27b
Qwen (Alibaba) 1440 1482 33.9 $0.80 1M tools 2026-08-14
Gemini 3.7 Flash
google/gemini-3.7-flash
Google 1491 1504 39.4 $1.50 1.048576M tools 2026-08-13
DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813
DeepSeek 1451 1471 36.3 $0.87 1.048576M tools 2026-08-12
Qwen3.8 2.4T A95B
qwen/qwen3.8-2.4t-a95b
Qwen (Alibaba) β€” β€” 40 $3.00 1.048576M tools 2026-08-12
Grok 4.6
x-ai/grok-4.6
xAI 1430 1465 44.4 $3.00 varies 500K tools 2026-08-12
DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
DeepSeek 1432 1456 34.5 $0.05 1.31072M tools 2026-07-31
Qwen3.7 Flash
qwen/qwen3.7-flash
Qwen (Alibaba) β€” β€” β€” $0.06 varies 1M tools 2026-07-27
Claude Opus 5
anthropic/claude-opus-5
Anthropic 1505 1532 50.7 $10.00 1M tools 2026-07-24
Gemini 3.6 Flash
google/gemini-3.6-flash
Google 1475 1491 34.3 $1.50 1.048576M tools 2026-07-21
Gemini 3.5 Flash Lite
google/gemini-3.5-flash-lite
Google 1436 1456 22.7 $0.85 1.048576M tools 2026-07-21
Kimi K3
moonshotai/kimi-k3
Moonshot AI 1472 1507 43.8 $4.61 1.048576M tools 2026-07-16
GPT-5.6 Luna Pro
openai/gpt-5.6-luna-pro
OpenAI β€” β€” β€” $0.45 varies 1.05M tools 2026-07-09
GPT-5.6 Luna
openai/gpt-5.6-luna
OpenAI 1431 1463 37.5 $0.45 varies 1.05M tools 2026-07-09
GPT-5.6 Terra Pro
openai/gpt-5.6-terra-pro
OpenAI β€” β€” β€” $4.50 varies 1.05M tools 2026-07-09
GPT-5.6 Terra
openai/gpt-5.6-terra
OpenAI 1446 1485 42.3 $4.50 varies 1.05M tools 2026-07-09
GPT-5.6 Sol Pro
openai/gpt-5.6-sol-pro
OpenAI β€” β€” β€” $4.00 varies 1.05M tools 2026-07-09
GPT-5.6 Sol
openai/gpt-5.6-sol
OpenAI 1454 1489 47.1 $4.00 varies 1.05M tools 2026-07-09
Grok 4.5
x-ai/grok-4.5
xAI 1451 1479 39.1 $3.00 varies 500K tools 2026-07-08
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
google/gemini-3.1-flash-lite-image
Google β€” β€” β€” $0.56 66K β€” 2026-06-30
Claude Sonnet 5
anthropic/claude-sonnet-5
Anthropic 1442 1485 38.4 $4.00 1M tools 2026-06-30
Nano Banana 2 (Gemini 3.1 Flash Image)
google/gemini-3.1-flash-image
Google β€” β€” β€” $1.13 131K β€” 2026-06-18
Nano Banana Pro (Gemini 3 Pro Image)
google/gemini-3-pro-image
Google β€” β€” β€” $4.50 131K tools 2026-06-18
GLM 5.2
z-ai/glm-5.2
Z.ai (Zhipu) 1467 1481 β€” $0.95 1.048576M tools 2026-06-16
Kimi K2.7 Code
moonshotai/kimi-k2.7-code
Moonshot AI β€” β€” 26.3 $1.41 262K tools 2026-06-12
Claude Fable 5
anthropic/claude-fable-5
Anthropic 1492 1519 49.7 $20.00 1M tools 2026-06-09
Qwen3.7 Plus
qwen/qwen3.7-plus
Qwen (Alibaba) 1454 1473 25.8 $0.56 varies 1M tools 2026-06-03
MiniMax M3
minimax/minimax-m3
MiniMax 1434 1471 29.6 $0.53 1.048576M tools 2026-05-31
Claude Opus 4.8
anthropic/claude-opus-4.8
Anthropic 1452 1487 42 $10.00 1M tools 2026-05-27
Qwen3.7 Max
qwen/qwen3.7-max
Qwen (Alibaba) 1474 1498 29.9 $2.21 1M tools 2026-05-21
Grok Build 0.1
x-ai/grok-build-0.1
xAI β€” β€” β€” $1.25 varies 256K tools 2026-05-20
Gemini 3.5 Flash
google/gemini-3.5-flash
Google 1483 1491 33 $3.38 1.048576M tools 2026-05-19
Grok 4.3
x-ai/grok-4.3
xAI 1398 1414 25.4 $1.56 varies 1M tools 2026-04-30
Mistral Medium 3.5
mistralai/mistral-medium-3-5
Mistral 1421 1461 14.9 $3.00 262K tools 2026-04-30
Qwen3.5 Plus 2026-04-20
qwen/qwen3.5-plus-20260420
Qwen (Alibaba) β€” β€” β€” $0.68 varies 1M tools 2026-04-27
DeepSeek V4 Pro 0423
deepseek/deepseek-v4-pro
DeepSeek 1451 1471 30.9 $1.02 1.048576M tools 2026-04-24
DeepSeek V4 Flash 0423
deepseek/deepseek-v4-flash
DeepSeek 1432 1456 24.8 $0.09 1.048576M tools 2026-04-24
Kimi K2.6
moonshotai/kimi-k2.6
Moonshot AI 1455 1486 β€” $1.71 262K tools 2026-04-20
Claude Opus 4.7
anthropic/claude-opus-4.7
Anthropic 1483 1516 β€” $10.00 1M tools 2026-04-16
GLM 5.1
z-ai/glm-5.1
Z.ai (Zhipu) 1462 1486 26.4 $1.49 205K tools 2026-04-07
GLM 5V Turbo
z-ai/glm-5v-turbo
Z.ai (Zhipu) 1436 1468 β€” $1.90 203K tools 2026-04-01
Grok 4.20 Multi-Agent
x-ai/grok-4.20-multi-agent
xAI β€” β€” β€” $1.56 varies 2M β€” 2026-03-31
Grok 4.20
x-ai/grok-4.20
xAI β€” β€” β€” $1.56 varies 2M tools 2026-03-31
MiniMax M2.7
minimax/minimax-m2.7
MiniMax 1404 1454 23.2 $0.53 205K tools 2026-03-18
Mistral Small 4
mistralai/mistral-small-2603
Mistral β€” β€” 11.5 $0.26 262K tools 2026-03-16
GLM 5 Turbo
z-ai/glm-5-turbo
Z.ai (Zhipu) β€” β€” β€” $1.90 203K tools 2026-03-15
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic 1458 1504 30.5 $6.00 1M tools 2026-02-17
MiniMax M2.5
minimax/minimax-m2.5
MiniMax 1359 1378 β€” $0.53 205K tools 2026-02-12
GLM 5
z-ai/glm-5
Z.ai (Zhipu) 1446 1461 β€” $0.93 205K tools 2026-02-11
Claude Opus 4.6
anthropic/claude-opus-4.6
Anthropic 1498 1534 β€” $10.00 1M tools 2026-02-04
Kimi K2.5
moonshotai/kimi-k2.5
Moonshot AI 1446 1485 β€” $0.90 262K tools 2026-01-27
MiniMax M2-her
minimax/minimax-m2-her
MiniMax β€” β€” β€” $0.53 66K β€” 2026-01-23
GLM 4.7 Flash
z-ai/glm-4.7-flash
Z.ai (Zhipu) 1352 1385 β€” $0.14 200K tools 2026-01-19
MiniMax M2.1
minimax/minimax-m2.1
MiniMax 1391 1422 β€” $0.53 205K tools 2025-12-23
Devstral 2 2512
mistralai/devstral-2512
Mistral β€” β€” 9.4 $0.80 262K tools 2025-12-09
Ministral 3 14B 2512
mistralai/ministral-14b-2512
Mistral β€” β€” 6 $0.20 262K tools 2025-12-02
Ministral 3 8B 2512
mistralai/ministral-8b-2512
Mistral β€” β€” 5.5 $0.15 262K tools 2025-12-02
Ministral 3 3B 2512
mistralai/ministral-3b-2512
Mistral β€” β€” 4.8 $0.10 131K tools 2025-12-02
DeepSeek V3.2
deepseek/deepseek-v3.2
DeepSeek 1425 1449 β€” $0.30 164K tools 2025-12-01
Mistral Large 3 2512
mistralai/mistral-large-2512
Mistral β€” β€” 9.7 $0.75 262K tools 2025-12-01
Kimi K2 Thinking
moonshotai/kimi-k2-thinking
Moonshot AI 1380 1399 β€” $1.07 262K tools 2025-11-06
Voxtral Small 24B 2507
mistralai/voxtral-small-24b-2507
Mistral β€” β€” β€” $0.15 33K tools 2025-10-30
MiniMax M2
minimax/minimax-m2
MiniMax 1342 1371 β€” $0.45 205K tools 2025-10-23

Pricing: OpenRouter model catalog, retrieved 2026-09-12. Elo capability: Arena leaderboard dataset (CC-BY-4.0), retrieved 2026-09-12. AA index: Artificial Analysis, republished in the OpenRouter catalog.

Top agent harnesses on SWE-bench Verified

Percentage of 500 real GitHub issues resolved. Entries are agent scaffold + model combinations, not bare models β€” the same model can move many points depending on the harness around it.

Read this board with care. We re-fetch it every day, but SWE-bench itself has not accepted a new Verified entry since 2026-02-26 β€” 198 days ago β€” so no 2026 frontier model appears here. Since late 2025 its submission policy has been limited to academic teams and established labs, which keeps most commercial agents off it entirely. Only 60 of 180 entries were re-run by the maintainers; the rest are submitter-reported.
#Agent + modelResolvedDateStatus
1Sonar Foundation Agent + Claude 4.5 Opus79.2%2025-12-05 self-reported
2live-SWE-agent + Claude 4.5 Opus medium (20251101)79.2%2025-12-15 self-reported
3TRAE + Doubao-Seed-Code78.8%2025-09-28 self-reported
4live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)77.4%2025-11-20 self-reported
5EPAM AI/Run Developer Agent v20250719 + Claude 4 Sonnet76.8%2025-08-04 self-reported
6Atlassian Rovo Dev (2025-09-02)76.8%2025-09-02 self-reported
7Claude 4.5 Opus (high)76.8%2026-02-17 self-reported
8ACoder76.4%2025-08-19 self-reported
9Gemini 3 Flash (high)75.8%2026-02-17 self-reported
10MiniMax M2.5 (high)75.8%2026-02-17 self-reported

Source: SWE-bench leaderboard Β· fetched 2026-09-12, board last updated 2026-02-26 Β· full top 20 + framework standings β†’

πŸ¦€ Spotlight: gcrab our project

A small, safe, replayable self-hosted agent β€” single static Go binary, zero deps, hard 8K context budget, policy-gated exec, append-only event log Β· 0 stars on GitHub.

gcrab.com β†’See it in the standings

Are you an agent?

This whole site is machine-readable with no key and open CORS. Ask it a question directly: GET /api/v1/route/coding.json returns the ranked capability-vs-cost table for that task. Also /llms.txt, /leaderboard.md, and the full API index. See API & Bots.