llm leaderboard

Best LLM for vision

Vision covers reading a document, following a diagram, and describing what is in a frame.

Rumeqo runs these models inside your team rooms. See what each one costs.

120 of 120 ranked models
Ranked models
rankmodelvendorcompositepricein / outbenchmarkseloeloeloeloeloelo
1
Claude Fable 5anthropic/claude-fable-5
Anthropic84.0$10.00 / $50.004 of 6 benchmarks1315133113631346
2
Qwen3.8 Maxqwen/qwen3.8-max
Qwen81.7$2.00 / $6.004 of 6 benchmarks1301131713271331
3
Claude Opus 4.6 (thinking)anthropic/claude-opus-4.6:thinking
Anthropic81.3$5.00 / $25.005 of 6 benchmarks13001315132513251269
4
Claude Opus 5 (high)anthropic/claude-opus-5:high
Anthropic81.0$5.00 / $25.004 of 6 benchmarks1297130813281344
5
Claude Opus 4.7anthropic/claude-opus-4.7
Anthropic80.9$5.00 / $25.004 of 6 benchmarks1299131413331327
6
Claude Opus 4.7 (thinking)anthropic/claude-opus-4.7:thinking
Anthropic80.5$5.00 / $25.005 of 6 benchmarks13011313133713301234
7
Gemini 3 Progoogle/gemini-3-pro
Google79.96 of 6 benchmarks128913031300130912991271
8
Grok 4.5x-ai/grok-4.5
SpaceXAI78.3$2.00 / $6.004 of 6 benchmarks1285129613281337
9
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
Google77.9$2.00 / $12.006 of 6 benchmarks127712941309131112861274
10
Claude Opus 4.8 (thinking)anthropic/claude-opus-4.8:thinking
Anthropic77.9$5.00 / $25.004 of 6 benchmarks1284129913201338
11
GPT-5.5openai/gpt-5.5
OpenAI77.6$5.00 / $30.004 of 6 benchmarks1286129913261330
12
GPT-5.5 (high)openai/gpt-5.5:high
OpenAI76.9$5.00 / $30.004 of 6 benchmarks1283129913141339
13
Claude Opus 4.6anthropic/claude-opus-4.6
Anthropic75.9$5.00 / $25.005 of 6 benchmarks12931310133113141222
14
Gemini 3.6 Flashgoogle/gemini-3.6-flash
Google75.1$1.50 / $7.503 of 6 benchmarks129513081313
15
Muse Sparkmeta/muse-spark
Meta74.44 of 6 benchmarks1294130213131300
16
Muse Spark 1.2 (xhigh)meta/muse-spark-1.2:xhigh
Meta74.4$1.25 / $4.253 of 6 benchmarks129012991316
17
Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium
Google72.8$1.50 / $9.004 of 6 benchmarks1284129612971322
18
GPT-5.4openai/gpt-5.4
OpenAI72.6$2.50 / $15.004 of 6 benchmarks1280129413181307
19
Gemini 3 Flash Previewgoogle/gemini-3-flash-preview
Google72.4$0.50 / $3.006 of 6 benchmarks127112851294130712931226
20
Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high
Anthropic72.0$2.00 / $10.004 of 6 benchmarks1273129113201313
21
Claude Opus 4.8anthropic/claude-opus-4.8
Anthropic71.6$5.00 / $25.004 of 6 benchmarks1280129313041322
22
GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh
OpenAI70.7$5.00 / $30.004 of 6 benchmarks1280128813081313
23
GPT-5.4 (high)openai/gpt-5.4:high
OpenAI70.1$2.50 / $15.005 of 6 benchmarks12831299131913301183
24
GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh
OpenAI69.8$1.00 / $6.004 of 6 benchmarks1270128013121317
25
Gemini 3.5 Flash (high)google/gemini-3.5-flash:high
Google69.6$1.50 / $9.004 of 6 benchmarks1283129213041294
26
Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal
Google68.7$0.50 / $3.006 of 6 benchmarks125912711284129812781213
27
GPT 5.5 Instantopenai/gpt-5.5-instant
OpenAI67.74 of 6 benchmarks1278128613151282
28
GPT-5.2 Chatopenai/gpt-5.2-chat
OpenAI67.3$1.75 / $14.005 of 6 benchmarks12781288131012931223
29
Muse Spark 1.1meta/muse-spark-1.1
Meta67.3$1.25 / $4.254 of 6 benchmarks1282129312991279
30
Kimi K2.6moonshotai/kimi-k2.6
MoonshotAI66.1$0.95 / $4.004 of 6 benchmarks1263127812921302
31
GPT-5.1 (high)openai/gpt-5.1:high
OpenAI64.7$1.25 / $10.006 of 6 benchmarks125012581275128212401240
32
Gemini 2.5 Progoogle/gemini-2.5-pro
Google63.8$1.25 / $10.006 of 6 benchmarks124612581266127212521251
33
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
Google63.2$0.30 / $2.503 of 6 benchmarks126012821283
34
Dola Seed 2.0 Probytedance/dola-seed-2.0-pro
ByteDance62.04 of 6 benchmarks1258127012871279
35
Qwen3.7 Plusqwen/qwen3.7-plus
Qwen61.6$0.32 / $1.284 of 6 benchmarks1262127812791276
36
Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning
xAI61.45 of 6 benchmarks12551264127312631247
37
Claude Sonnet 4.6anthropic/claude-sonnet-4.6
Anthropic61.3$3.00 / $15.005 of 6 benchmarks12751291130613011159
38
Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking
MoonshotAI60.8$0.57 / $2.856 of 6 benchmarks124912641277129012301196
39
Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b
Qwen60.2$0.50 / $3.605 of 6 benchmarks12471262127612831226
40
Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview
Google60.0$0.25 / $1.506 of 6 benchmarks123412471251127012721226
41
GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh
OpenAI59.9$0.10 / $0.604 of 6 benchmarks1249125612811285
42
GPT-5.2 (high)openai/gpt-5.2:high
OpenAI59.7$1.75 / $14.006 of 6 benchmarks124412601274128411851268
43
Grok 4.20 Multi Agent Beta 0309x-ai/grok-4.20-multi-agent-beta-0309
xAI59.65 of 6 benchmarks12521261126812731231
44
GPT-5.4 Mini (high)openai/gpt-5.4-mini:high
OpenAI57.1$0.75 / $4.505 of 6 benchmarks12531266128412931182
45
Grok 4.3x-ai/grok-4.3
SpaceXAI55.6$1.25 / $2.505 of 6 benchmarks12441253126812561227
46
Gemma 4 31Bgoogle/gemma-4-31b-it
Google55.3$0.10 / $0.346 of 6 benchmarks125612701281130011911165
47
Gemma 4 26B A4B google/gemma-4-26b-a4b-it
Google54.6$0.12 / $0.406 of 6 benchmarks124012561261128511531248
48
Chatgpt 4oopenai/chatgpt-4o
OpenAI54.56 of 6 benchmarks124112491264124712251212
49
MiniMax M3minimax/minimax-m3
MiniMax53.7$0.30 / $1.204 of 6 benchmarks1240125312631263
50
GPT 4.5 Previewopenai/gpt-4.5-preview
OpenAI52.51 of 6 benchmarks1226
51
GPT-5 (high)openai/gpt-5:high
OpenAI52.4$1.25 / $10.006 of 6 benchmarks121312321252126012571191
52
GPT 5 Chatopenai/gpt-5-chat
OpenAI50.66 of 6 benchmarks122512451260127411881198
53
Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant
Moonshot AI49.04 of 6 benchmarks1237124512491247
54
MiMo-V2.5xiaomi/mimo-v2.5
Xiaomi48.8$0.14 / $0.285 of 6 benchmarks12381253127112601172
55
Gemini 2.5 Flashgoogle/gemini-2.5-flash
Google48.4$0.30 / $2.506 of 6 benchmarks121412221233123812201215
56
Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b
Qwen48.1$0.29 / $2.404 of 6 benchmarks1227123912461252
57
GPT-5.1openai/gpt-5.1
OpenAI47.9$1.25 / $10.006 of 6 benchmarks123812521264125411881180
58
Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct
Qwen47.7$0.26 / $1.046 of 6 benchmarks121512291239126011901208
59
o1openai/o1
OpenAI47.2$15.00 / $60.001 of 6 benchmarks1193
60
o3openai/o3
OpenAI46.8$2.00 / $8.006 of 6 benchmarks121712231231124712341181
61
GPT-5.2openai/gpt-5.2
OpenAI46.5$1.75 / $14.006 of 6 benchmarks122912381257126711921167
62
Mimo v2 Omnixiaomi/mimo-v2-omni
Xiaomi46.54 of 6 benchmarks1217123112541246
63
GLM 5V Turboz-ai/glm-5v-turbo
Z.ai46.4$1.20 / $4.005 of 6 benchmarks12311244126412541181
64
Gemini 1.5 Pro 002google/gemini-1.5-pro-002
Google44.91 of 6 benchmarks1180
65
GPT-4.1openai/gpt-4.1
OpenAI44.2$2.00 / $8.006 of 6 benchmarks121412261232124911831207
66
GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13
OpenAI43.5$5.00 / $15.001 of 6 benchmarks1162
67
Ernie 5.0 Preview 1220baidu/ernie-5.0-preview-1220
Baidu43.44 of 6 benchmarks1218123012231237
68
Claude Sonnet 4 (thinking 32K)anthropic/claude-sonnet-4:thinking-32k
Anthropic43.0$3.00 / $15.003 of 6 benchmarks120812161228
69
o4 Miniopenai/o4-mini
OpenAI42.5$1.10 / $4.406 of 6 benchmarks120212101220125112181182
70
Qwen VL (max)qwen/qwen-vl:max
Qwen42.03 of 6 benchmarks118612141249
71
Qwen3.5-27Bqwen/qwen3.5-27b
Qwen40.8$0.195 / $1.565 of 6 benchmarks12191232124212421149
72
Mistral Large 3mistralai/mistral-large-3
Mistral AI40.44 of 6 benchmarks1199121912411221
73
GPT-4.1 Miniopenai/gpt-4.1-mini
OpenAI40.2$0.40 / $1.606 of 6 benchmarks120312061209122211981195
74
Gemini 1.5 Flash 002google/gemini-1.5-flash-002
Google40.21 of 6 benchmarks1141
75
Mistral Medium 3.5mistralai/mistral-medium-3-5
Mistral39.8$1.50 / $7.504 of 6 benchmarks1198121512421216
76
Gemini 2.0 Flash Lite Preview 02 05google/gemini-2.0-flash-lite-preview-02-05
Google39.61 of 6 benchmarks1136
77
Qwen2.5 VL 72B Instructqwen/qwen2.5-vl-72b-instruct
Qwen38.51 of 6 benchmarks1122
78
Gemini 1.5 Pro 001google/gemini-1.5-pro-001
Google38.21 of 6 benchmarks1120
79
Claude Opus 4 (thinking 16K)anthropic/claude-opus-4:thinking-16k
Anthropic38.1$15.00 / $75.004 of 6 benchmarks1207121612031216
80
Qwen2.5 VL 32B Instructqwen/qwen2.5-vl-32b-instruct
Qwen37.91 of 6 benchmarks1119
81
GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06
OpenAI37.6$2.50 / $10.001 of 6 benchmarks1119
82
GPT-5 Mini (high)openai/gpt-5-mini:high
OpenAI37.6$0.25 / $2.006 of 6 benchmarks118311981214123312071168
83
GPT-4 Turboopenai/gpt-4-turbo
OpenAI37.4$10.00 / $30.001 of 6 benchmarks1112
84
Claude 3.7 Sonnet (thinking 32K)anthropic/claude-3-7-sonnet:thinking-32k
Anthropic37.14 of 6 benchmarks1196121012161210
85
GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18
OpenAI36.8$0.15 / $0.601 of 6 benchmarks1098
86
Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking
Qwen36.7$0.40 / $4.004 of 6 benchmarks1190120112101227
87
Claude Opus 4anthropic/claude-opus-4
Anthropic36.7$15.00 / $75.004 of 6 benchmarks1189119712031237
88
GPT-4.1 Nanoopenai/gpt-4.1-nano
OpenAI36.5$0.10 / $0.401 of 6 benchmarks1089
89
Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning
xAI36.45 of 6 benchmarks11951192121511701202
90
Gemini 1.5 Flash 8B 001google/gemini-1.5-flash-8b-001
Google36.31 of 6 benchmarks1072
91
Claude 3 Opusanthropic/claude-3-opus
Anthropic36.01 of 6 benchmarks1062
92
Gemini 1.5 Flash 001google/gemini-1.5-flash-001
Google35.71 of 6 benchmarks1060
93
Grok 4 0709x-ai/grok-4-0709
xAI35.66 of 6 benchmarks118211751191116912361167
94
Amazon Nova Pro v1.0amazon/amazon-nova-pro-v1.0
Amazon35.41 of 6 benchmarks1019
95
Amazon Nova Lite v1.0amazon/amazon-nova-lite-v1.0
Amazon35.11 of 6 benchmarks1019
96
Claude Sonnet 4anthropic/claude-sonnet-4
Anthropic35.1$3.00 / $15.004 of 6 benchmarks1189119012021222
97
Claude 3 Sonnetanthropic/claude-3-sonnet
Anthropic34.81 of 6 benchmarks1017
98
GPT-5.4 Nano (high)openai/gpt-5.4-nano:high
OpenAI34.7$0.20 / $1.255 of 6 benchmarks12021215122912351119
99
Claude 3 Haikuanthropic/claude-3-haiku
Anthropic34.6$0.25 / $1.251 of 6 benchmarks1001
100
Gemini 2.5 Flash Lite (thinking)google/gemini-2.5-flash-lite:thinking
Google33.9$0.10 / $0.406 of 6 benchmarks118811881187119011851185
101
Claude 3.7 Sonnetanthropic/claude-3-7-sonnet
Anthropic33.94 of 6 benchmarks1176118611921228
102
Hunyuan Vision 1.5 (thinking)tencent/hunyuan-vision-1.5:thinking
Tencent31.34 of 6 benchmarks1159116111861229
103
Gemini 2.5 Flash Lite (nothinking)google/gemini-2.5-flash-lite:nothinking
Google31.2$0.10 / $0.404 of 6 benchmarks1174118511831194
104
Claude 3.5 Sonnetanthropic/claude-3-5-sonnet
Anthropic29.24 of 6 benchmarks1161117711841170
105
GLM 4.6Vz-ai/glm-4.6v
Z.ai27.8$0.30 / $0.904 of 6 benchmarks1164117111861155
106
Gemini 2.0 Flash 001google/gemini-2.0-flash-001
Google27.75 of 6 benchmarks11721166117611901148
107
GPT-5 Nano (high)openai/gpt-5-nano:high
OpenAI26.1$0.05 / $0.404 of 6 benchmarks1148115811531187
108
Hunyuan Large Visiontencent/hunyuan-large-vision
Tencent26.03 of 6 benchmarks115011461114
109
GLM 4.5Vz-ai/glm-4.5v
Z.ai25.9$0.60 / $1.804 of 6 benchmarks1156115811691170
110
Step 3stepfun/step-3
StepFun24.44 of 6 benchmarks1145115111461176
111
Gemma 3 27Bgoogle/gemma-3-27b-it
Google24.0$0.08 / $0.456 of 6 benchmarks116011621166117211751118
112
Step 1o Turbo 202506stepfun/step-1o-turbo-202506
StepFun23.94 of 6 benchmarks1157115411621138
113
Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct
Meta23.64 of 6 benchmarks1147115311521163
114
Mistral Medium 2508mistralai/mistral-medium-2508
Mistral AI23.06 of 6 benchmarks115911721178117011321130
115
Molmo 2 8Ballenai/molmo-2-8b
Allen Institute for AI22.33 of 6 benchmarks110810871108
116
Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct
Meta21.14 of 6 benchmarks1128113011481149
117
Claude 3.5 Haikuanthropic/claude-3-5-haiku
Anthropic20.74 of 6 benchmarks1128112911351150
118
Mistral Medium 2505mistralai/mistral-medium-2505
Mistral AI20.06 of 6 benchmarks115611591167116511231092
119
Mistral Small 2506mistralai/mistral-small-2506
Mistral AI19.06 of 6 benchmarks114111421140117611341055
120
Mistral Small 3.1 24B Instruct 2503mistralai/mistral-small-3.1-24b-instruct-2503
Mistral AI17.96 of 6 benchmarks112811331130115911271135
How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 3 of the 6 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

Data sources

Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.

  • LMArenaCC BY 4.0

    Arena ratings by LMArena, from the public leaderboard dataset.