llm leaderboard

Best AI for design

Design here means producing an interface a person prefers, not writing about design.

Rumeqo runs these models inside your team rooms. See what each one costs.

115 of 115 ranked models
Ranked models
rankmodelvendorcompositepricein / outbenchmarkseloeloeloelo
1
Claude Opus 5 (max)anthropic/claude-opus-5:max
Anthropic82.4$5.00 / $25.004 of 4 benchmarks1691171116791670
2
Qwen3.8 Maxqwen/qwen3.8-max
Qwen80.6$2.00 / $6.004 of 4 benchmarks1669167916281631
3
Claude Fable 5anthropic/claude-fable-5
Anthropic80.0$10.00 / $50.004 of 4 benchmarks1627163016711626
4
Kimi K3 (max)moonshotai/kimi-k3:max
MoonshotAI79.2$3.00 / $15.004 of 4 benchmarks1674169216501570
5
GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh
OpenAI78.6$5.00 / $30.004 of 4 benchmarks1622162916231581
6
Claude Opus 5 (high)anthropic/claude-opus-5:high
Anthropic77.6$5.00 / $25.003 of 4 benchmarks166416811661
7
Grok 4.5x-ai/grok-4.5
SpaceXAI75.9$2.00 / $6.004 of 4 benchmarks1553155515781580
8
Grok 4.6 (high)x-ai/grok-4.6:high
SpaceXAI75.8$2.00 / $6.003 of 4 benchmarks161816321619
9
Claude Opus 4.7anthropic/claude-opus-4.7
Anthropic74.7$5.00 / $25.004 of 4 benchmarks1558155415621565
10
Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high
Anthropic73.5$5.00 / $25.003 of 4 benchmarks156415621557
11
GLM 5.2 (max)z-ai/glm-5.2:max
Z.ai73.2$0.63 / $1.983 of 4 benchmarks158715971540
12
Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high
Anthropic72.8$5.00 / $25.003 of 4 benchmarks155715581554
13
DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high
DeepSeek72.1$0.14 / $0.283 of 4 benchmarks158215921538
14
Claude Opus 4.6 (high)anthropic/claude-opus-4.6:high
Anthropic71.7$5.00 / $25.003 of 4 benchmarks154515381559
15
Muse Spark 1.1meta/muse-spark-1.1
Meta71.1$1.25 / $4.254 of 4 benchmarks1539153015391544
16
Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high
Anthropic70.0$2.00 / $10.004 of 4 benchmarks1541154615231540
17
Claude Opus 4.6anthropic/claude-opus-4.6
Anthropic69.8$5.00 / $25.004 of 4 benchmarks1537152915501538
18
Claude Opus 4.8anthropic/claude-opus-4.8
Anthropic68.4$5.00 / $25.003 of 4 benchmarks153915421525
19
Muse Spark 1.2 (xhigh)meta/muse-spark-1.2:xhigh
Meta68.1$1.25 / $4.253 of 4 benchmarks153515301539
20
Gemini 3.6 Flash (high)google/gemini-3.6-flash:high
Google67.7$1.50 / $7.503 of 4 benchmarks153715301533
21
GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh
OpenAI66.8$1.00 / $6.004 of 4 benchmarks1523151415421525
22
Claude Sonnet 4.6anthropic/claude-sonnet-4.6
Anthropic66.7$3.00 / $15.004 of 4 benchmarks1524151315291530
23
GPT-5.5 (xhigh)openai/gpt-5.5:xhigh
OpenAI65.5$5.00 / $30.004 of 4 benchmarks1509150015441526
24
Qwen3.7 Maxqwen/qwen3.7-max
Qwen64.5$1.475 / $4.4253 of 4 benchmarks151715151521
25
Hy3tencent/hy3
Tencent64.1$0.132 / $0.5283 of 4 benchmarks152315361464
26
GLM 5.1z-ai/glm-5.1
Z.ai63.7$1.40 / $4.403 of 4 benchmarks151115061521
27
Seed 2.1 Pro Previewbytedance/seed-2.1-pro-preview
ByteDance62.74 of 4 benchmarks1522152414971506
28
GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh
OpenAI62.6$0.10 / $0.604 of 4 benchmarks1518150715371493
29
Kimi K2.6moonshotai/kimi-k2.6
MoonshotAI62.0$0.95 / $4.004 of 4 benchmarks1509150414951520
30
Gemini 3.5 Flash (high)google/gemini-3.5-flash:high
Google61.6$1.50 / $9.004 of 4 benchmarks1506149815411492
31
GPT-5.5 (high)openai/gpt-5.5:high
OpenAI61.6$5.00 / $30.004 of 4 benchmarks1486147715391511
32
Claude Opus 4.5 (high 32K)anthropic/claude-opus-4.5:high-32k
Anthropic61.3$5.00 / $25.003 of 4 benchmarks149414801512
33
Claude Opus 4.7 (thinking)anthropic/claude-opus-4.7:thinking
Anthropic60.8$5.00 / $25.001 of 4 benchmarks1579
34
Qwen3.6 Max Previewqwen/qwen3.6-max-preview
Qwen59.4$1.027 / $6.1623 of 4 benchmarks147914701495
35
Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium
Google59.1$1.50 / $9.004 of 4 benchmarks1488148215261490
36
MiMo-V2.5-Proxiaomi/mimo-v2.5-pro
Xiaomi58.3$0.435 / $0.873 of 4 benchmarks147414621482
37
Claude Opus 4.5anthropic/claude-opus-4.5
Anthropic57.9$5.00 / $25.003 of 4 benchmarks146814571482
38
Claude Opus 4.6 (thinking)anthropic/claude-opus-4.6:thinking
Anthropic57.5$5.00 / $25.001 of 4 benchmarks1540
39
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
MoonshotAI57.4$0.67 / $3.404 of 4 benchmarks1473147314621512
40
DeepSeek V4 Pro (high)deepseek/deepseek-v4-pro:high
DeepSeek56.7$1.168 / $2.3363 of 4 benchmarks146414551473
41
GPT-5.4 (high)openai/gpt-5.4:high
OpenAI56.5$2.50 / $15.003 of 4 benchmarks146314531477
42
GPT-5.5openai/gpt-5.5
OpenAI56.2$5.00 / $30.004 of 4 benchmarks1458144615041499
43
MiniMax M3minimax/minimax-m3
MiniMax55.8$0.30 / $1.204 of 4 benchmarks1491148714641475
44
Gemini 3.6 Flashgoogle/gemini-3.6-flash
Google54.3$1.50 / $7.501 of 4 benchmarks1529
45
GPT-5.4 (medium)openai/gpt-5.4:medium
OpenAI54.2$2.50 / $15.003 of 4 benchmarks144214291474
46
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
Google53.3$2.00 / $12.004 of 4 benchmarks1447143714941481
47
DeepSeek V4 Prodeepseek/deepseek-v4-pro
DeepSeek53.0$1.168 / $2.3363 of 4 benchmarks144514391436
48
Qwen3.6 Plusqwen/qwen3.6-plus
Qwen52.1$0.325 / $1.954 of 4 benchmarks1459145414611470
49
GLM 4.7z-ai/glm-4.7
Z.ai51.2$0.40 / $1.753 of 4 benchmarks143414291446
50
GLM 5z-ai/glm-5
Z.ai50.7$0.95 / $2.553 of 4 benchmarks143614251454
51
MiMo-V2.5xiaomi/mimo-v2.5
Xiaomi50.4$0.14 / $0.283 of 4 benchmarks143814271428
52
Mimo v2 Proxiaomi/mimo-v2-pro
Xiaomi50.23 of 4 benchmarks143414261439
53
GPT-5 (medium)openai/gpt-5:medium
OpenAI49.9$1.25 / $10.002 of 4 benchmarks14191429
54
Gemini 3 Flash Previewgoogle/gemini-3-flash-preview
Google49.5$0.50 / $3.004 of 4 benchmarks1438143014441455
55
GPT-5.2openai/gpt-5.2
OpenAI49.4$1.75 / $14.002 of 4 benchmarks14181428
56
Gemini 3 Progoogle/gemini-3-pro
Google47.74 of 4 benchmarks1438139614691448
57
Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking
MoonshotAI46.6$0.57 / $2.854 of 4 benchmarks1436142914291441
58
Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b
Qwen45.6$0.50 / $3.603 of 4 benchmarks140013891413
59
GPT-5.1 (medium)openai/gpt-5.1:medium
OpenAI45.3$1.25 / $10.002 of 4 benchmarks13911401
60
GPT-5.4 Mini (high)openai/gpt-5.4-mini:high
OpenAI45.2$0.75 / $4.503 of 4 benchmarks139713871423
61
Inklingthinkingmachines/inkling
Thinking Machines44.8$0.95 / $4.053 of 4 benchmarks140513991381
62
MiniMax M2.7minimax/minimax-m2.7
MiniMax44.6$0.30 / $1.203 of 4 benchmarks139813881402
63
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
Google44.5$0.30 / $2.503 of 4 benchmarks144914261426
64
GPT-5.3-Codexopenai/gpt-5.3-codex
OpenAI44.2$1.75 / $14.004 of 4 benchmarks1409139814281446
65
Claude Opus 4.1anthropic/claude-opus-4.1
Anthropic44.2$15.00 / $75.002 of 4 benchmarks13891398
66
Claude Sonnet 4.5 (high 32K)anthropic/claude-sonnet-4.5:high-32k
Anthropic43.3$3.00 / $15.003 of 4 benchmarks139213861399
67
GPT-5.4openai/gpt-5.4
OpenAI42.7$2.50 / $15.004 of 4 benchmarks1390138814161449
68
MiniMax M2.5minimax/minimax-m2.5
MiniMax42.5$0.22 / $0.903 of 4 benchmarks138413731415
69
MiniMax M2.1minimax/minimax-m2.1
MiniMax41.7$0.30 / $1.203 of 4 benchmarks138713691401
70
Claude Sonnet 4.5anthropic/claude-sonnet-4.5
Anthropic41.0$3.00 / $15.003 of 4 benchmarks138613841390
71
Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant
Moonshot AI40.54 of 4 benchmarks1405139214141418
72
Solar Pro 4upstage/solar-pro4
Upstage39.6$0.03 / $0.123 of 4 benchmarks137213581394
73
Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning
xAI38.63 of 4 benchmarks137413671367
74
Gemma 4 31Bgoogle/gemma-4-31b-it
Google38.3$0.10 / $0.343 of 4 benchmarks136513581378
75
Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal
Google37.8$0.50 / $3.004 of 4 benchmarks1383137413981432
76
Gemma 4 26B A4B google/gemma-4-26b-a4b-it
Google37.7$0.12 / $0.403 of 4 benchmarks136213541372
77
Muse Glimmer 30Bmeta/muse-glimmer-30b
Meta37.5$0.35 / $1.503 of 4 benchmarks135913521385
78
GLM 5V Turboz-ai/glm-5v-turbo
Z.ai37.4$1.20 / $4.004 of 4 benchmarks1400138513651423
79
DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking
DeepSeek36.7$0.269 / $0.403 of 4 benchmarks136113501369
80
Qwen3.5-27Bqwen/qwen3.5-27b
Qwen36.6$0.195 / $1.563 of 4 benchmarks135713451393
81
GLM 4.6z-ai/glm-4.6
Z.ai36.4$0.50 / $2.002 of 4 benchmarks13401350
82
GPT-5.1 (high)openai/gpt-5.1:high
OpenAI36.4$1.25 / $10.001 of 4 benchmarks1420
83
GPT-5.1-Codexopenai/gpt-5.1-codex
OpenAI35.5$1.25 / $10.002 of 4 benchmarks13361346
84
Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b
Qwen35.0$0.29 / $2.403 of 4 benchmarks135813501357
85
Hunyuan Hy3 Previewtencent/hunyuan-hy3-preview
Tencent34.73 of 4 benchmarks135613481360
86
MiniMax M2minimax/minimax-m2
MiniMax32.8$0.255 / $1.022 of 4 benchmarks12971307
87
Laguna M.1poolside/laguna-m.1
Poolside32.43 of 4 benchmarks134813391324
88
GPT-5.2-Codexopenai/gpt-5.2-codex
OpenAI32.2$1.75 / $14.003 of 4 benchmarks133813301348
89
Mimo v2 Flashxiaomi/mimo-v2-flash
Xiaomi31.53 of 4 benchmarks133013101353
90
Grok 4.3x-ai/grok-4.3
SpaceXAI31.2$1.25 / $2.504 of 4 benchmarks1355135213601375
91
DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp
DeepSeek30.9$0.27 / $0.412 of 4 benchmarks12721282
92
Claude Haiku 4.5anthropic/claude-haiku-4.5
Anthropic30.3$1.00 / $5.003 of 4 benchmarks132613231316
93
Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo
Moonshot AI30.33 of 4 benchmarks132313171327
94
DeepSeek V3.2deepseek/deepseek-v3.2
DeepSeek30.2$0.269 / $0.403 of 4 benchmarks132413351304
95
KAT-Coder-Pro V1kwaipilot/kat-coder-pro-v1
Kwaipilot29.82 of 4 benchmarks12551265
96
GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini
OpenAI28.9$0.25 / $2.002 of 4 benchmarks12441254
97
GPT-5.1openai/gpt-5.1
OpenAI28.1$1.25 / $10.004 of 4 benchmarks1341130513641345
98
Mimo v2 Flash (thinking)xiaomi/mimo-v2-flash:thinking
Xiaomi28.03 of 4 benchmarks129312311346
99
Mistral Medium 3.5mistralai/mistral-medium-3-5
Mistral27.8$1.50 / $7.503 of 4 benchmarks126612551317
100
Laguna XS.2poolside/laguna-xs.2
Poolside27.63 of 4 benchmarks130312961281
101
Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct
Qwen27.23 of 4 benchmarks127312571286
102
Grok 4.1 (thinking)x-ai/grok-4.1:thinking
xAI26.42 of 4 benchmarks12101220
103
Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b
Qwen26.1$0.25 / $1.253 of 4 benchmarks125012341298
104
Grok Code Fast 1x-ai/grok-code-fast-1
xAI24.52 of 4 benchmarks11641172
105
Qwen3.5-Flashqwen/qwen3.5-flash-02-23
Qwen24.5$0.065 / $0.263 of 4 benchmarks123812241293
106
Grok 4 Fast (reasoning)x-ai/grok-4-fast:reasoning
xAI24.12 of 4 benchmarks11611170
107
Trinity Large Thinkingarcee-ai/trinity-large-thinking
Arcee AI23.6$0.22 / $0.853 of 4 benchmarks123912331241
108
Devstral Medium 2507mistralai/devstral-medium-2507
Mistral AI23.62 of 4 benchmarks10801087
109
Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning
xAI23.43 of 4 benchmarks124012221252
110
Mistral Large 3mistralai/mistral-large-3
Mistral AI23.33 of 4 benchmarks123012391361
111
Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview
Google21.7$0.25 / $1.504 of 4 benchmarks1254124212801336
112
Gemini 2.5 Progoogle/gemini-2.5-pro
Google21.4$1.25 / $10.003 of 4 benchmarks122612351283
113
Granite 4.1 8Bibm-granite/granite-4.1-8b
IBM21.0$0.05 / $0.103 of 4 benchmarks119211791219
114
Devstral 2mistralai/devstral-2
Mistral AI20.53 of 4 benchmarks119411391216
115
Mercury 2inception/mercury-2
Inception20.2$0.25 / $0.753 of 4 benchmarks116611561199
How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 2 of the 4 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

Data sources

Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.

  • LMArenaCC BY 4.0

    Arena ratings by LMArena, from the public leaderboard dataset.