llm leaderboard

Best LLM for coding

Ranked on issue resolution, webdev preference and security review. Every row prints its coverage, and a model these benchmarks have only partly reached still ranks, on the results it does have, under a partial coverage mark.

Rumeqo runs these models inside your team rooms. See what each one costs.

317 of 317 ranked models
Ranked models
rankmodelvendorcompositepricein / outbenchmarks%%%%%%%eloelo%count
1
Claude Opus 5 (max)anthropic/claude-opus-5:max
Anthropic81.0$5.00 / $25.003 of 11 benchmarks50.0%15301691
2
Claude Opus 4.6anthropic/claude-opus-4.6
Anthropic79.5$5.00 / $25.004 of 11 benchmarks78.7%72.0%15471537
3
Kimi K3 (max)moonshotai/kimi-k3:max
MoonshotAI78.0$3.00 / $15.004 of 11 benchmarks1543167437.2%44.0
4
Claude Opus 4.5anthropic/claude-opus-4.5
Anthropic76.8$5.00 / $25.005 of 11 benchmarks76.7%52.6%70.7%15231468
5
Claude Fable 5anthropic/claude-fable-5
Anthropic76.4$10.00 / $50.002 of 11 benchmarks15541627
6
Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high
Anthropic75.8$5.00 / $25.003 of 11 benchmarks31.1%15521557
7
Claude Opus 5 (high)anthropic/claude-opus-5:high
Anthropic75.7$5.00 / $25.002 of 11 benchmarks15311664
8
Qwen3.8 Maxqwen/qwen3.8-max
Qwen75.6$2.00 / $6.002 of 11 benchmarks15291669
9
Claude Opus 4.7anthropic/claude-opus-4.7
Anthropic74.7$5.00 / $25.002 of 11 benchmarks15471558
10
GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh
OpenAI74.7$5.00 / $30.002 of 11 benchmarks15261622
11
GPT-5.6 Luna (high)openai/gpt-5.6-luna:high
OpenAI73.5$0.10 / $0.602 of 11 benchmarks41.9%57.0
12
Muse Spark 1.1meta/muse-spark-1.1
Meta73.1$1.25 / $4.252 of 11 benchmarks15311539
13
Muse Spark 1.2 (xhigh)meta/muse-spark-1.2:xhigh
Meta72.4$1.25 / $4.252 of 11 benchmarks15331535
14
Grok 4.5x-ai/grok-4.5
SpaceXAI72.2$2.00 / $6.002 of 11 benchmarks15211553
15
Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high
Anthropic72.2$2.00 / $10.002 of 11 benchmarks15231541
16
Grok 4.6 (high)x-ai/grok-4.6:high
SpaceXAI71.9$2.00 / $6.002 of 11 benchmarks15121618
17
Claude Opus 4.8anthropic/claude-opus-4.8
Anthropic71.9$5.00 / $25.002 of 11 benchmarks15241539
18
Gemini 3 Flash Previewgoogle/gemini-3-flash-preview
Google71.6$0.50 / $3.004 of 11 benchmarks75.4%72.7%15081438
19
Gemini 3.6 Flash (high)google/gemini-3.6-flash:high
Google71.4$1.50 / $7.502 of 11 benchmarks15221537
20
GLM 5.1z-ai/glm-5.1
Z.ai70.0$1.40 / $4.403 of 11 benchmarks74.2%15151511
21
GPT-5.5 (high)openai/gpt-5.5:high
OpenAI69.8$5.00 / $30.005 of 11 benchmarks10.0%1520148647.7%72.0
22
Claude Sonnet 4.6anthropic/claude-sonnet-4.6
Anthropic69.8$3.00 / $15.005 of 11 benchmarks75.2%1528152429.1%32.0
23
Claude Fable 5 (high)anthropic/claude-fable-5:high
Anthropic69.7$10.00 / $50.001 of 11 benchmarks63.9%
24
Qwen3.6 Max Previewqwen/qwen3.6-max-preview
Qwen69.5$1.027 / $6.1623 of 11 benchmarks76.7%15091479
25
GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh
OpenAI69.4$1.00 / $6.002 of 11 benchmarks15161523
26
Claude Opus 4.5 (high 32K)anthropic/claude-opus-4.5:high-32k
Anthropic69.3$5.00 / $25.002 of 11 benchmarks15301494
27
GPT 5.5 Pre Release (xhigh)openai/gpt-5.5-pre-release:xhigh
OpenAI69.31 of 11 benchmarks80.6%
28
Claude Opus 4.5 (medium)anthropic/claude-opus-4.5:medium
Anthropic68.6$5.00 / $25.001 of 11 benchmarks79.2%
29
Doubao-Seed-Codebytedance/doubao-seed-code
ByteDance68.21 of 11 benchmarks78.8%
30
Claude Fable 5 (max)anthropic/claude-fable-5:max
Anthropic67.8$10.00 / $50.001 of 11 benchmarks39.5%
31
Grok 4.5 (high)x-ai/grok-4.5:high
SpaceXAI67.7$2.00 / $6.002 of 11 benchmarks38.4%41.0
32
Muse Sparkmeta/muse-spark
Meta67.71 of 11 benchmarks1526
33
GLM 5.2 (max)z-ai/glm-5.2:max
Z.ai67.1$0.63 / $1.984 of 11 benchmarks78.7%9.5%15061587
34
DeepSeek V4 Pro (max)deepseek/deepseek-v4-pro:max
DeepSeek67.1$1.168 / $2.3361 of 11 benchmarks77.6%
35
Hy3tencent/hy3
Tencent67.1$0.132 / $0.5282 of 11 benchmarks15031523
36
Claude Opus 4.7 (max)anthropic/claude-opus-4.7:max
Anthropic67.0$5.00 / $25.002 of 11 benchmarks83.5%19.1%
37
DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high
DeepSeek66.9$0.14 / $0.282 of 11 benchmarks14801582
38
MiMo-V2.5-Proxiaomi/mimo-v2.5-pro
Xiaomi66.9$0.435 / $0.872 of 11 benchmarks15201474
39
GPT-5.5 (xhigh)openai/gpt-5.5:xhigh
OpenAI66.5$5.00 / $30.002 of 11 benchmarks34.3%1509
40
Qwen3.7 Maxqwen/qwen3.7-max
Qwen66.2$1.475 / $4.4254 of 11 benchmarks77.3%9.5%15251517
41
Claude Opus 4.5 (high)anthropic/claude-opus-4.5:high
Anthropic66.0$5.00 / $25.001 of 11 benchmarks76.8%
42
GPT-5.6 Sol (max)openai/gpt-5.6-sol:max
OpenAI65.8$5.00 / $30.001 of 11 benchmarks39.0%
43
GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh
OpenAI65.8$0.10 / $0.602 of 11 benchmarks14991518
44
GPT-5.4 (high)openai/gpt-5.4:high
OpenAI65.7$2.50 / $15.004 of 11 benchmarks76.9%15.6%15211463
45
GPT-5.2 Chatopenai/gpt-5.2-chat
OpenAI65.6$1.75 / $14.001 of 11 benchmarks1515
46
Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium
Google65.3$1.50 / $9.002 of 11 benchmarks15071488
47
Ernie 5.1baidu/ernie-5.1
Baidu65.31 of 11 benchmarks1514
48
GPT 5.5 Instantopenai/gpt-5.5-instant
OpenAI65.31 of 11 benchmarks1514
49
Qwen3.5 Max Previewqwen/qwen3.5-max-preview
Qwen64.91 of 11 benchmarks1513
50
Dola Seed 2.0 Probytedance/dola-seed-2.0-pro
ByteDance64.71 of 11 benchmarks1513
51
Claude Opus 4.1 (thinking 16K)anthropic/claude-opus-4.1:thinking-16k
Anthropic64.5$15.00 / $75.001 of 11 benchmarks1512
52
Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high
Google64.2$0.50 / $3.001 of 11 benchmarks75.8%
53
MiniMax M2.5 (high)minimax/minimax-m2.5:high
MiniMax64.2$0.22 / $0.901 of 11 benchmarks75.8%
54
Gemini 3 Progoogle/gemini-3-pro
Google64.23 of 11 benchmarks68.7%15181438
55
MiniMax M3minimax/minimax-m3
MiniMax64.1$0.30 / $1.202 of 11 benchmarks14971491
56
GPT-5.5openai/gpt-5.5
OpenAI63.8$5.00 / $30.002 of 11 benchmarks15101458
57
Grok 4.20 Multi Agent Beta 0309x-ai/grok-4.20-multi-agent-beta-0309
xAI63.71 of 11 benchmarks1508
58
Gemini 3.1 Pro Preview Custom Toolsgoogle/gemini-3.1-pro-preview-customtools
Google63.7$2.00 / $12.001 of 11 benchmarks75.6%
59
GLM 5z-ai/glm-5
Z.ai63.4$0.95 / $2.554 of 11 benchmarks72.1%69.7%14971436
60
Qwen3.7 Plusqwen/qwen3.7-plus
Qwen63.3$0.32 / $1.281 of 11 benchmarks1506
61
Claude Sonnet 4anthropic/claude-sonnet-4
Anthropic63.2$3.00 / $15.004 of 11 benchmarks57.0%58.3%35.6%1449
62
Seed 2.1 Pro Previewbytedance/seed-2.1-pro-preview
ByteDance62.81 of 11 benchmarks1522
63
GPT-5.3-Codex (high)openai/gpt-5.3-codex:high
OpenAI62.5$1.75 / $14.001 of 11 benchmarks74.8%
64
Gemini 3.5 Flash (high)google/gemini-3.5-flash:high
Google62.4$1.50 / $9.004 of 11 benchmarks79.3%4.8%15091506
65
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
Google62.4$0.30 / $2.502 of 11 benchmarks15031449
66
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
Google62.3$2.00 / $12.003 of 11 benchmarks14.3%15211447
67
Claude Opus 4.6 (high)anthropic/claude-opus-4.6:high
Anthropic62.3$5.00 / $25.004 of 11 benchmarks1552154526.7%24.0
68
GPT-5openai/gpt-5
OpenAI62.2$1.25 / $10.001 of 11 benchmarks74.4%
69
Claude Opus 4 (thinking 16K)anthropic/claude-opus-4:thinking-16k
Anthropic62.1$15.00 / $75.001 of 11 benchmarks1499
70
GPT-5.5 (low)openai/gpt-5.5:low
OpenAI61.9$5.00 / $30.002 of 11 benchmarks32.6%38.0
71
Claude Opus 4.8 (max)anthropic/claude-opus-4.8:max
Anthropic61.9$5.00 / $25.001 of 11 benchmarks28.6%
72
DeepSeek V4 Prodeepseek/deepseek-v4-pro
DeepSeek61.7$1.168 / $2.3362 of 11 benchmarks15021445
73
GPT-5.3 Chatopenai/gpt-5.3-chat
OpenAI61.11 of 11 benchmarks1496
74
DeepSeek V4 Pro (high)deepseek/deepseek-v4-pro:high
DeepSeek61.1$1.168 / $2.3362 of 11 benchmarks14891464
75
Grok 4.1x-ai/grok-4.1
xAI60.71 of 11 benchmarks1492
76
o3openai/o3
OpenAI60.6$2.00 / $8.003 of 11 benchmarks58.4%36.0%1460
77
Claude Sonnet 4.5 (high 32K)anthropic/claude-sonnet-4.5:high-32k
Anthropic60.5$3.00 / $15.002 of 11 benchmarks15191392
78
Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking
MoonshotAI60.5$0.57 / $2.852 of 11 benchmarks15021436
79
Mimo v2 Proxiaomi/mimo-v2-pro
Xiaomi60.22 of 11 benchmarks15031434
80
Kimi K2.6moonshotai/kimi-k2.6
MoonshotAI60.2$0.95 / $4.004 of 11 benchmarks76.7%2.4%15141509
81
GPT-5.4 (xhigh)openai/gpt-5.4:xhigh
OpenAI59.9$2.50 / $15.001 of 11 benchmarks25.4%
82
Gemini 3 Pro Previewgoogle/gemini-3-pro-preview
Google59.91 of 11 benchmarks72.9%
83
Ernie 5.0 0110baidu/ernie-5.0-0110
Baidu59.71 of 11 benchmarks1490
84
Claude 3.7 Sonnetanthropic/claude-3-7-sonnet
Anthropic59.65 of 11 benchmarks61.0%51.7%33.8%31.3%1430
85
MiMo-V2.5xiaomi/mimo-v2.5
Xiaomi59.6$0.14 / $0.282 of 11 benchmarks14911438
86
GLM 5 (high)z-ai/glm-5:high
Z.ai59.3$0.95 / $2.551 of 11 benchmarks72.8%
87
Mimo v2 Omnixiaomi/mimo-v2-omni
Xiaomi59.31 of 11 benchmarks1486
88
Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high
Anthropic59.3$5.00 / $25.004 of 11 benchmarks1533156424.4%24.0
89
GPT-5.4openai/gpt-5.4
OpenAI59.3$2.50 / $15.002 of 11 benchmarks15141390
90
Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant
Moonshot AI59.22 of 11 benchmarks15051405
91
DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash
DeepSeek58.9$0.14 / $0.281 of 11 benchmarks1483
92
GPT-5.1 (high)openai/gpt-5.1:high
OpenAI58.8$1.25 / $10.002 of 11 benchmarks68.0%1491
93
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
MoonshotAI58.8$0.67 / $3.401 of 11 benchmarks1473
94
Kimi K2.5moonshotai/kimi-k2.5
MoonshotAI58.4$0.57 / $2.852 of 11 benchmarks73.8%67.3%
95
Claude Sonnet 4.5 (high)anthropic/claude-sonnet-4.5:high
Anthropic58.0$3.00 / $15.001 of 11 benchmarks71.4%
96
GPT-5.2 (xhigh)openai/gpt-5.2:xhigh
OpenAI58.0$1.75 / $14.001 of 11 benchmarks23.0%
97
Inklingthinkingmachines/inkling
Thinking Machines57.9$0.95 / $4.052 of 11 benchmarks14941405
98
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b
NVIDIA57.8$0.60 / $3.601 of 11 benchmarks1475
99
GLM 4.7z-ai/glm-4.7
Z.ai57.7$0.40 / $1.752 of 11 benchmarks14851434
100
Kimi K2 0905moonshotai/kimi-k2-0905
MoonshotAI57.6$0.60 / $2.502 of 11 benchmarks71.2%1468
101
DeepSeek V3.2 Exp (thinking)deepseek/deepseek-v3.2-exp:thinking
DeepSeek57.5$0.27 / $0.411 of 11 benchmarks1475
102
GPT-5.2 (high)openai/gpt-5.2:high
OpenAI57.5$1.75 / $14.003 of 11 benchmarks73.8%66.7%1490
103
Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning
xAI57.32 of 11 benchmarks15111374
104
LongCat Flash Chatmeituan/longcat-flash-chat
Meituan57.31 of 11 benchmarks1474
105
Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b
Qwen57.2$0.50 / $3.602 of 11 benchmarks14911400
106
Claude Sonnet 4 (thinking 32K)anthropic/claude-sonnet-4:thinking-32k
Anthropic57.1$3.00 / $15.001 of 11 benchmarks1473
107
GPT-5.4 Mini (high)openai/gpt-5.4-mini:high
OpenAI57.1$0.75 / $4.502 of 11 benchmarks14971397
108
Qwen3.6 Plusqwen/qwen3.6-plus
Qwen57.0$0.325 / $1.953 of 11 benchmarks57.9%14951459
109
Qwen3 Maxqwen/qwen3-max
Qwen57.0$0.78 / $3.901 of 11 benchmarks1473
110
Kimi K2.5 (high)moonshotai/kimi-k2.5:high
MoonshotAI56.9$0.57 / $2.851 of 11 benchmarks70.8%
111
Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507
Qwen56.8$0.09 / $0.551 of 11 benchmarks1472
112
GPT-5 (medium)openai/gpt-5:medium
OpenAI56.7$1.25 / $10.002 of 11 benchmarks71.5%1419
113
Ernie 5.0 Preview 1203baidu/ernie-5.0-preview-1203
Baidu56.71 of 11 benchmarks1472
114
GLM 5V Turboz-ai/glm-5v-turbo
Z.ai56.6$1.20 / $4.002 of 11 benchmarks14901400
115
GPT-5.2openai/gpt-5.2
OpenAI56.6$1.75 / $14.003 of 11 benchmarks69.0%14821418
116
Claude Opus 4anthropic/claude-opus-4
Anthropic56.5$15.00 / $75.002 of 11 benchmarks70.7%1464
117
GPT-5.6 Sol (high)openai/gpt-5.6-sol:high
OpenAI56.4$5.00 / $30.001 of 11 benchmarks20.0%
118
Chatgpt 4oopenai/chatgpt-4o
OpenAI56.31 of 11 benchmarks1468
119
DeepSeek V3.2 (high)deepseek/deepseek-v3.2:high
DeepSeek56.1$0.269 / $0.401 of 11 benchmarks70.0%
120
GPT-5.4 (medium)openai/gpt-5.4:medium
OpenAI56.1$2.50 / $15.001 of 11 benchmarks1442
121
GPT-5 (high)openai/gpt-5:high
OpenAI56.0$1.25 / $10.003 of 11 benchmarks73.5%12.7%1469
122
Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct
Qwen55.6$0.26 / $1.041 of 11 benchmarks1465
123
Gemini 3 Pro Preview (high)google/gemini-3-pro-preview:high
Google55.51 of 11 benchmarks69.6%
124
R1 0528deepseek/deepseek-r1-0528
DeepSeek55.5$0.50 / $2.151 of 11 benchmarks1464
125
DeepSeek V3.1 Terminus (thinking)deepseek/deepseek-v3.1-terminus:thinking
DeepSeek55.2$0.27 / $0.951 of 11 benchmarks1463
126
GPT 5 Chatopenai/gpt-5-chat
OpenAI55.11 of 11 benchmarks1462
127
MiniMax M2.7minimax/minimax-m2.7
MiniMax55.0$0.30 / $1.202 of 11 benchmarks14791398
128
Gemma 4 31Bgoogle/gemma-4-31b-it
Google55.0$0.10 / $0.342 of 11 benchmarks14991365
129
GPT-5.4 Nano (high)openai/gpt-5.4-nano:high
OpenAI54.5$0.20 / $1.251 of 11 benchmarks1460
130
Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal
Google54.4$0.50 / $3.002 of 11 benchmarks14911383
131
Claude Haiku 4.5 (high)anthropic/claude-haiku-4.5:high
Anthropic54.2$1.00 / $5.001 of 11 benchmarks66.6%
132
GPT 4.5 Previewopenai/gpt-4.5-preview
OpenAI54.11 of 11 benchmarks1459
133
Claude Sonnet 4.5anthropic/claude-sonnet-4.5
Anthropic54.0$3.00 / $15.006 of 11 benchmarks71.3%44.3%67.0%2.4%15131386
134
GPT-5.1-Codex (medium)openai/gpt-5.1-codex:medium
OpenAI53.6$1.25 / $10.001 of 11 benchmarks66.0%
135
DeepSeek V3.1 (thinking)deepseek/deepseek-chat-v3.1:thinking
DeepSeek53.5$0.25 / $0.951 of 11 benchmarks1457
136
Kimi K2 0711moonshotai/kimi-k2
MoonshotAI53.5$0.57 / $2.302 of 11 benchmarks65.4%1461
137
Mistral Medium 2508mistralai/mistral-medium-2508
Mistral AI53.31 of 11 benchmarks1455
138
DeepSeek V4 Pro (xhigh)deepseek/deepseek-v4-pro:xhigh
DeepSeek53.3$1.168 / $2.3362 of 11 benchmarks26.7%30.0
139
Claude Opus 4.1anthropic/claude-opus-4.1
Anthropic53.2$15.00 / $75.004 of 11 benchmarks73.3%7.9%15051389
140
Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking
Qwen53.1$0.40 / $4.001 of 11 benchmarks1455
141
Claude Opus 4.5 (128K)anthropic/claude-opus-4.5:128k
Anthropic53.1$5.00 / $25.001 of 11 benchmarks14.3%
142
GPT-5.3-Codexopenai/gpt-5.3-codex
OpenAI53.1$1.75 / $14.001 of 11 benchmarks1409
143
Claude 3.7 Sonnet (thinking 32K)anthropic/claude-3-7-sonnet:thinking-32k
Anthropic52.91 of 11 benchmarks1452
144
Step 3.5 Flashstepfun/step-3.5-flash
StepFun52.7$0.10 / $0.301 of 11 benchmarks1451
145
GPT-5 Mini (medium)openai/gpt-5-mini:medium
OpenAI52.7$0.25 / $2.001 of 11 benchmarks64.7%
146
Claude 3.5 Haikuanthropic/claude-3-5-haiku
Anthropic52.42 of 11 benchmarks41.7%1385
147
Gemma 4 26B A4B google/gemma-4-26b-a4b-it
Google52.3$0.12 / $0.402 of 11 benchmarks14811362
148
Claude 3.5 Sonnetanthropic/claude-3-5-sonnet
Anthropic52.35 of 11 benchmarks62.8%51.3%24.9%25.3%1435
149
DeepSeek V3.1deepseek/deepseek-chat-v3.1
DeepSeek52.2$0.25 / $0.951 of 11 benchmarks1448
150
Kimi K2 Thinkingmoonshotai/kimi-k2-thinking
MoonshotAI51.9$0.60 / $2.501 of 11 benchmarks63.4%
151
Muse Glimmer 30Bmeta/muse-glimmer-30b
Meta51.9$0.35 / $1.502 of 11 benchmarks14811359
152
Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct
Qwen51.9$0.10 / $1.101 of 11 benchmarks1446
153
Qwen3 235B A22B (nothinking)qwen/qwen3-235b-a22b:nothinking
Qwen51.8$0.455 / $1.821 of 11 benchmarks1446
154
Grok 4.3x-ai/grok-4.3
SpaceXAI51.6$1.25 / $2.502 of 11 benchmarks14881355
155
R1deepseek/deepseek-r1
DeepSeek51.6$0.70 / $2.501 of 11 benchmarks1445
156
Grok 3 Betax-ai/grok-3-beta
xAI51.31 of 11 benchmarks1443
157
Trinity Large Previewarcee-ai/trinity-large-preview
Arcee AI51.21 of 11 benchmarks1443
158
o3 (medium)openai/o3:medium
OpenAI51.2$2.00 / $8.001 of 11 benchmarks62.3%
159
Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507
Qwen51.1$0.23 / $2.301 of 11 benchmarks1442
160
GPT-5.1 (medium)openai/gpt-5.1:medium
OpenAI50.9$1.25 / $10.002 of 11 benchmarks66.0%1391
161
Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507
Qwen50.8$0.0482 / $0.19311 of 11 benchmarks1440
162
DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus
DeepSeek50.7$0.27 / $0.951 of 11 benchmarks1439
163
Hunyuan Vision 1.5 (thinking)tencent/hunyuan-vision-1.5:thinking
Tencent50.51 of 11 benchmarks1438
164
MiniMax M2.5minimax/minimax-m2.5
MiniMax50.2$0.22 / $0.903 of 11 benchmarks68.3%14441384
165
Grok 4 0709x-ai/grok-4-0709
xAI49.81 of 11 benchmarks1435
166
o3 Mini Highopenai/o3-mini-high
OpenAI49.7$1.10 / $4.401 of 11 benchmarks1435
167
GPT-5.1openai/gpt-5.1
OpenAI49.6$1.25 / $10.002 of 11 benchmarks14741341
168
Mistral Medium 2505mistralai/mistral-medium-2505
Mistral AI49.41 of 11 benchmarks1433
169
Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo
Moonshot AI49.42 of 11 benchmarks14861323
170
DeepSeek V3deepseek/deepseek-chat
DeepSeek49.3$0.2574 / $1.02872 of 11 benchmarks36.7%1388
171
DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking
DeepSeek49.3$0.269 / $0.403 of 11 benchmarks60.0%14751361
172
Qwen3 235B A22Bqwen/qwen3-235b-a22b
Qwen49.2$0.455 / $1.821 of 11 benchmarks1433
173
Claude Opus 4.6 (max)anthropic/claude-opus-4.6:max
Anthropic49.1$5.00 / $25.001 of 11 benchmarks12.7%
174
o4 Miniopenai/o4-mini
OpenAI48.9$1.10 / $4.403 of 11 benchmarks45.0%33.9%1433
175
o1openai/o1
OpenAI48.9$15.00 / $60.002 of 11 benchmarks64.6%1433
176
Ernie 5.0 Preview 1022baidu/ernie-5.0-preview-1022
Baidu48.91 of 11 benchmarks1432
177
GPT-5 Mini (high)openai/gpt-5-mini:high
OpenAI48.6$0.25 / $2.001 of 11 benchmarks1431
178
Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b
Qwen48.4$0.29 / $2.402 of 11 benchmarks14591358
179
Hunyuan Hy3 Previewtencent/hunyuan-hy3-preview
Tencent48.42 of 11 benchmarks14611356
180
DeepSeek V3 0324deepseek/deepseek-chat-v3-0324
DeepSeek48.3$0.27 / $1.121 of 11 benchmarks1429
181
Solar Pro 4upstage/solar-pro4
Upstage48.3$0.03 / $0.122 of 11 benchmarks14501372
182
MiniMax M2.1minimax/minimax-m2.1
MiniMax48.2$0.30 / $1.202 of 11 benchmarks14401387
183
GLM 4.5 Airz-ai/glm-4.5-air
Z.ai48.2$0.13 / $0.851 of 11 benchmarks1426
184
Devstral Small 2512mistralai/devstral-small-2512
Mistral AI48.11 of 11 benchmarks56.4%
185
GLM 4.7 Flashz-ai/glm-4.7-flash
Z.ai48.1$0.06 / $0.401 of 11 benchmarks1424
186
Grok 4.1 (thinking)x-ai/grok-4.1:thinking
xAI47.82 of 11 benchmarks14991210
187
Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning-30b-a3b
NVIDIA47.81 of 11 benchmarks1422
188
GLM 4.5z-ai/glm-4.5
Z.ai47.7$0.60 / $2.202 of 11 benchmarks54.2%1455
189
Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking
Qwen47.6$0.15 / $1.201 of 11 benchmarks1421
190
GLM 4.6Vz-ai/glm-4.6v
Z.ai47.5$0.30 / $0.901 of 11 benchmarks1417
191
Claude Sonnet 5anthropic/claude-sonnet-5
Anthropic47.5$2.00 / $10.002 of 11 benchmarks25.6%27.0
192
MiniMax M1minimax/minimax-m1
MiniMax47.2$0.55 / $2.201 of 11 benchmarks1416
193
Mistral Medium 3.5mistralai/mistral-medium-3-5
Mistral47.1$1.50 / $7.502 of 11 benchmarks14791266
194
Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct
Qwen47.03 of 11 benchmarks69.6%14571273
195
Mistral Small 2506mistralai/mistral-small-2506
Mistral AI47.01 of 11 benchmarks1412
196
Qwen3.5-27Bqwen/qwen3.5-27b
Qwen46.8$0.195 / $1.562 of 11 benchmarks14501357
197
Ling Flash 2.0inclusionai/ling-flash-2.0
inclusionAI46.81 of 11 benchmarks1411
198
INTELLECT-3prime-intellect/intellect-3
Prime Intellect46.71 of 11 benchmarks1409
199
Step 3stepfun/step-3
StepFun46.61 of 11 benchmarks1408
200
Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b
NVIDIA46.4$0.085 / $0.401 of 11 benchmarks1408
201
GPT-4.1openai/gpt-4.1
OpenAI46.3$2.00 / $8.003 of 11 benchmarks48.5%31.1%1456
202
Qwen3 32Bqwen/qwen3-32b
Qwen46.3$0.08 / $0.281 of 11 benchmarks1407
203
Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct
Qwen46.3$0.07 / $0.281 of 11 benchmarks51.6%
204
GLM 4.5Vz-ai/glm-4.5v
Z.ai46.1$0.60 / $1.801 of 11 benchmarks1405
205
Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5
NVIDIA46.01 of 11 benchmarks1404
206
Qwen 2.5 (max)qwen/qwen-2.5:max
Qwen45.91 of 11 benchmarks1403
207
Hunyuan T1tencent/hunyuan-t1
Tencent45.71 of 11 benchmarks1399
208
DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp
DeepSeek45.7$0.27 / $0.412 of 11 benchmarks14651272
209
Gemini 2.5 Flash Lite (nothinking)google/gemini-2.5-flash-lite:nothinking
Google45.6$0.10 / $0.401 of 11 benchmarks1397
210
GPT-5.2-Codexopenai/gpt-5.2-codex
OpenAI45.5$1.75 / $14.003 of 11 benchmarks72.8%66.3%1338
211
Devstral Small 2505mistralai/devstral-small-2505
Mistral AI45.51 of 11 benchmarks46.8%
212
Laguna M.1poolside/laguna-m.1
Poolside45.51 of 11 benchmarks1348
213
Nova 2 Liteamazon/nova-2-lite-v1
Amazon45.5$0.30 / $2.501 of 11 benchmarks1395
214
Hunyuan TurboStencent/hunyuan-turbos
Tencent45.31 of 11 benchmarks1394
215
Llama 3.1 Nemotron Ultra 253B v1nvidia/llama-3.1-nemotron-ultra-253b-v1
NVIDIA45.01 of 11 benchmarks1391
216
Ring Flash 2.0inclusionai/ring-flash-2.0
inclusionAI44.91 of 11 benchmarks1390
217
Kimi K2 Instructmoonshotai/kimi-k2-instruct
Moonshot AI44.71 of 11 benchmarks43.8%
218
Mimo v2 Flashxiaomi/mimo-v2-flash
Xiaomi44.72 of 11 benchmarks14461330
219
Grok 3 Mini Beta (high)x-ai/grok-3-mini-beta:high
xAI44.61 of 11 benchmarks1390
220
o3 Miniopenai/o3-mini
OpenAI44.5$1.10 / $4.403 of 11 benchmarks42.4%32.3%1416
221
Command A (03-2025)cohere/command-a-03-2025
Cohere44.51 of 11 benchmarks1390
222
GPT-5.1-Codexopenai/gpt-5.1-codex
OpenAI44.3$1.25 / $10.001 of 11 benchmarks1336
223
Magistral Medium 2506mistralai/magistral-medium-2506
Mistral AI44.21 of 11 benchmarks1387
224
Nova Premier 1.0amazon/nova-premier-v1
Amazon44.2$2.50 / $12.501 of 11 benchmarks42.4%
225
O1 Miniopenai/o1-mini
OpenAI44.11 of 11 benchmarks1387
226
GLM 4.6z-ai/glm-4.6
Z.ai44.1$0.50 / $2.003 of 11 benchmarks55.4%14581340
227
Grok 3 Mini Betax-ai/grok-3-mini-beta
xAI43.91 of 11 benchmarks1386
228
Mistral Large 3mistralai/mistral-large-3
Mistral AI43.92 of 11 benchmarks14681230
229
Qwen3 30B A3Bqwen/qwen3-30b-a3b
Qwen43.8$0.12 / $0.501 of 11 benchmarks1386
230
Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning
xAI43.72 of 11 benchmarks14611240
231
Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview
Google43.4$0.25 / $1.502 of 11 benchmarks14571254
232
QwQ 32Bqwen/qwq-32b
Qwen43.41 of 11 benchmarks1384
233
GPT-5 Nano (high)openai/gpt-5-nano:high
OpenAI43.3$0.05 / $0.401 of 11 benchmarks1384
234
Gemini 2.5 Flash Lite (thinking)google/gemini-2.5-flash-lite:thinking
Google43.1$0.10 / $0.401 of 11 benchmarks1384
235
OLMo 3.1 32B Instructallenai/olmo-3.1-32b-instruct
Allen Institute for AI43.01 of 11 benchmarks1382
236
GPT-4.1 Nanoopenai/gpt-4.1-nano
OpenAI42.9$0.10 / $0.401 of 11 benchmarks1374
237
Laguna XS.2poolside/laguna-xs.2
Poolside42.81 of 11 benchmarks1303
238
DeepSeek V4 Flash 0423 (xhigh)deepseek/deepseek-v4-flash:xhigh
DeepSeek42.7$0.14 / $0.282 of 11 benchmarks20.9%27.0
239
gpt-oss-20bopenai/gpt-oss-20b
OpenAI42.6$0.03 / $0.131 of 11 benchmarks1370
240
Claude Haiku 4.5anthropic/claude-haiku-4.5
Anthropic42.5$1.00 / $5.003 of 11 benchmarks64.7%14791326
241
Devstral Small 2507mistralai/devstral-small-2507
Mistral AI42.51 of 11 benchmarks38.0%
242
Gemini 2.5 Progoogle/gemini-2.5-pro
Google42.3$1.25 / $10.003 of 11 benchmarks57.6%14651226
243
Mercuryinception/mercury
Inception42.31 of 11 benchmarks1367
244
GPT-5 Nano (medium)openai/gpt-5-nano:medium
OpenAI42.1$0.05 / $0.401 of 11 benchmarks34.8%
245
OLMo 3 32B Thinkallenai/olmo-3-32b-think
Allen Institute for AI42.01 of 11 benchmarks1364
246
Llama 3.3 Nemotron Super 49B v1nvidia/llama-3.3-nemotron-super-49b-v1
NVIDIA41.91 of 11 benchmarks1363
247
Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b
NVIDIA41.8$0.05 / $0.201 of 11 benchmarks1362
248
GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20
OpenAI41.7$2.50 / $10.001 of 11 benchmarks31.0%
249
Mistral Small 3.1 24B Instruct 2503mistralai/mistral-small-3.1-24b-instruct-2503
Mistral AI41.51 of 11 benchmarks1362
250
Gemma 3 27Bgoogle/gemma-3-27b-it
Google41.2$0.08 / $0.451 of 11 benchmarks1358
251
GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13
OpenAI41.1$5.00 / $15.004 of 11 benchmarks38.8%31.3%12.0%1369
252
Gemini 1.5 Pro 002google/gemini-1.5-pro-002
Google41.11 of 11 benchmarks1356
253
KAT-Coder-Pro V1kwaipilot/kat-coder-pro-v1
Kwaipilot40.91 of 11 benchmarks1255
254
Hunyuan Large Visiontencent/hunyuan-large-vision
Tencent40.91 of 11 benchmarks1356
255
Mimo v2 Flash (thinking)xiaomi/mimo-v2-flash:thinking
Xiaomi40.92 of 11 benchmarks14311293
256
Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct
Qwen40.8$0.36 / $0.401 of 11 benchmarks1356
257
Mistral Large 2407mistralai/mistral-large-2407
Mistral40.5$2.00 / $6.001 of 11 benchmarks1354
258
Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b
Qwen40.4$0.25 / $1.252 of 11 benchmarks14351250
259
Step 1o Turbo 202506stepfun/step-1o-turbo-202506
StepFun40.21 of 11 benchmarks1352
260
GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18
OpenAI40.1$0.15 / $0.601 of 11 benchmarks1349
261
GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini
OpenAI40.0$0.25 / $2.001 of 11 benchmarks1244
262
GPT-4.1 Miniopenai/gpt-4.1-mini
OpenAI40.0$0.40 / $1.602 of 11 benchmarks23.9%1433
263
GPT-4 Turboopenai/gpt-4-turbo
OpenAI40.0$10.00 / $30.001 of 11 benchmarks1347
264
Qwen3.5-Flashqwen/qwen3.5-flash-02-23
Qwen39.8$0.065 / $0.262 of 11 benchmarks14371238
265
Gemini 1.5 Pro 001google/gemini-1.5-pro-001
Google39.81 of 11 benchmarks1347
266
DeepSeek V3.2deepseek/deepseek-v3.2
DeepSeek39.8$0.269 / $0.403 of 11 benchmarks59.0%14701324
267
Mistral Large 2411mistralai/mistral-large-2411
Mistral AI39.71 of 11 benchmarks1346
268
Gemini 2.5 Flashgoogle/gemini-2.5-flash
Google39.6$0.30 / $2.502 of 11 benchmarks28.7%1424
269
GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06
OpenAI39.6$2.50 / $10.004 of 11 benchmarks27.0%39.7%30.4%1360
270
Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct
Meta39.6$0.10 / $0.321 of 11 benchmarks1346
271
Qwen 2.5qwen/qwen-2.5
Qwen39.62 of 11 benchmarks40.2%24.7%
272
Amazon Nova Pro v1.0amazon/amazon-nova-pro-v1.0
Amazon39.41 of 11 benchmarks1343
273
Gemini 2.0 Flash Lite Preview 02 05google/gemini-2.0-flash-lite-preview-02-05
Google39.31 of 11 benchmarks1343
274
OLMo 3.1 32B Thinkallenai/olmo-3.1-32b-think
Allen Institute for AI38.91 of 11 benchmarks1338
275
Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct
Meta38.7$0.40 / $0.401 of 11 benchmarks1333
276
Claude 3 Sonnetanthropic/claude-3-sonnet
Anthropic38.61 of 11 benchmarks1318
277
Gemma 3 12Bgoogle/gemma-3-12b-it
Google38.5$0.05 / $0.151 of 11 benchmarks1316
278
MiniMax M2minimax/minimax-m2
MiniMax38.4$0.255 / $1.023 of 11 benchmarks61.0%13851297
279
GPT 4 1106 Previewopenai/gpt-4-1106-preview
OpenAI38.44 of 11 benchmarks22.4%28.3%12.5%1340
280
Gemini 1.5 Flash 002google/gemini-1.5-flash-002
Google38.31 of 11 benchmarks1316
281
Mistral Small 24B Instruct 2501mistralai/mistral-small-24b-instruct-2501
Mistral AI38.21 of 11 benchmarks1312
282
Gemini 1.5 Flash 001google/gemini-1.5-flash-001
Google38.01 of 11 benchmarks1309
283
Grok 4 Fast (reasoning)x-ai/grok-4-fast:reasoning
xAI37.92 of 11 benchmarks14361161
284
Gemma 3n E4Bgoogle/gemma-3n-e4b-it
Google37.91 of 11 benchmarks1308
285
Amazon Nova Lite v1.0amazon/amazon-nova-lite-v1.0
Amazon37.81 of 11 benchmarks1306
286
Trinity Large Thinkingarcee-ai/trinity-large-thinking
Arcee AI37.6$0.22 / $0.852 of 11 benchmarks14141239
287
Amazon Nova Micro v1.0amazon/amazon-nova-micro-v1.0
Amazon37.51 of 11 benchmarks1289
288
Command R (08-2024)cohere/command-r-08-2024
Cohere37.4$0.15 / $0.601 of 11 benchmarks1281
289
Command R+ (08-2024)cohere/command-r-plus-08-2024
Cohere37.2$2.50 / $10.001 of 11 benchmarks1280
290
OLMo 2 0325 32B Instructallenai/olmo-2-0325-32b-instruct
Allen Institute for AI37.11 of 11 benchmarks1280
291
Grok Code Fast 1x-ai/grok-code-fast-1
xAI37.01 of 11 benchmarks1164
292
Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct
Mistral37.0$2.00 / $6.001 of 11 benchmarks1277
293
Gemma 3 4Bgoogle/gemma-3-4b-it
Google36.81 of 11 benchmarks1274
294
gpt-oss-120bopenai/gpt-oss-120b
OpenAI36.7$0.03 / $0.172 of 11 benchmarks26.0%1390
295
Gemini 1.5 Flash 8B 001google/gemini-1.5-flash-8b-001
Google36.71 of 11 benchmarks1272
296
Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct
Meta36.5$0.05 / $0.081 of 11 benchmarks1260
297
Gemini 3.1 Pro Preview (high)google/gemini-3.1-pro-preview:high
Google36.4$2.00 / $12.001 of 11 benchmarks8.9%
298
Devstral Medium 2507mistralai/devstral-medium-2507
Mistral AI36.41 of 11 benchmarks1080
299
GPT-4oopenai/gpt-4o
OpenAI36.4$2.50 / $10.001 of 11 benchmarks12.2%
300
QwQ 32B Previewqwen/qwq-32b-preview
Qwen36.41 of 11 benchmarks1173
301
Devstral 2mistralai/devstral-2
Mistral AI36.12 of 11 benchmarks53.8%1194
302
Claude Opus 4.8 (medium)anthropic/claude-opus-4.8:medium
Anthropic35.9$5.00 / $25.002 of 11 benchmarks20.9%19.0
303
GPT-5 Miniopenai/gpt-5-mini
OpenAI35.8$0.25 / $2.002 of 11 benchmarks56.2%39.7%
304
Mercury 2inception/mercury-2
Inception34.6$0.25 / $0.752 of 11 benchmarks13941166
305
Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct
Meta34.32 of 11 benchmarks21.0%1373
306
Claude 3 Opusanthropic/claude-3-opus
Anthropic34.34 of 11 benchmarks15.8%26.3%10.5%1355
307
Claude 3 Haikuanthropic/claude-3-haiku
Anthropic33.6$0.25 / $1.252 of 11 benchmarks40.6%1301
308
Gemini 2.0 Flash 001google/gemini-2.0-flash-001
Google33.32 of 11 benchmarks13.5%1365
309
Claude 2anthropic/claude-2
Anthropic32.83 of 11 benchmarks4.4%3.0%2.0%
310
Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct
Meta32.62 of 11 benchmarks9.1%1362
311
Granite 4.1 8Bibm-granite/granite-4.1-8b
IBM31.2$0.05 / $0.102 of 11 benchmarks13531192
312
GLM 5.2 (high)z-ai/glm-5.2:high
Z.ai31.1$0.63 / $1.982 of 11 benchmarks17.4%18.0
313
Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct
Qwen30.52 of 11 benchmarks9.0%1342
314
SWE Llamaprinceton-nlp/swe-llama
Princeton NLP28.13 of 11 benchmarks1.4%1.3%0.7%
315
Claude Opus 4.7 (medium)anthropic/claude-opus-4.7:medium
Anthropic27.3$5.00 / $25.002 of 11 benchmarks7.0%7.0
316
SWE Llama 13Bprinceton-nlp/swe-llama-13b
Princeton NLP26.53 of 11 benchmarks1.2%1.0%0.7%
317
GPT 3.5openai/gpt-3.5
OpenAI21.83 of 11 benchmarks0.4%0.3%0.2%
How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 3 of the 11 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

Data sources

Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.

  • Epoch AI Benchmarking HubCC BY 4.0

    Benchmark runs by Epoch AI, from the AI Benchmarking Hub.

  • LMArenaCC BY 4.0

    Arena ratings by LMArena, from the public leaderboard dataset.

  • SWE-bench

    Resolve rates published by the SWE-bench maintainers.

  • Warden

    Security review results published by Warden.