Best LLM for vision
Vision covers reading a document, following a diagram, and describing what is in a frame.
Rumeqo runs these models inside your team rooms. See what each one costs.
| rank | model | vendor | composite | pricein / out | benchmarks | elo | elo | elo | elo | elo | elo |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5anthropic/claude-fable-5 | Anthropic | 84.0 | $10.00 / $50.00 | 4 of 6 benchmarks | 1315 | 1331 | 1363 | 1346 | ||
| 2 | Qwen3.8 Maxqwen/qwen3.8-max | Qwen | 81.7 | $2.00 / $6.00 | 4 of 6 benchmarks | 1301 | 1317 | 1327 | 1331 | ||
| 3 | Claude Opus 4.6 (thinking)anthropic/claude-opus-4.6:thinking | Anthropic | 81.3 | $5.00 / $25.00 | 5 of 6 benchmarks | 1300 | 1315 | 1325 | 1325 | 1269 | |
| 4 | Claude Opus 5 (high)anthropic/claude-opus-5:high | Anthropic | 81.0 | $5.00 / $25.00 | 4 of 6 benchmarks | 1297 | 1308 | 1328 | 1344 | ||
| 5 | Claude Opus 4.7anthropic/claude-opus-4.7 | Anthropic | 80.9 | $5.00 / $25.00 | 4 of 6 benchmarks | 1299 | 1314 | 1333 | 1327 | ||
| 6 | Claude Opus 4.7 (thinking)anthropic/claude-opus-4.7:thinking | Anthropic | 80.5 | $5.00 / $25.00 | 5 of 6 benchmarks | 1301 | 1313 | 1337 | 1330 | 1234 | |
| 7 | Gemini 3 Progoogle/gemini-3-pro | 79.9 | — | 6 of 6 benchmarks | 1289 | 1303 | 1300 | 1309 | 1299 | 1271 | |
| 8 | Grok 4.5x-ai/grok-4.5 | SpaceXAI | 78.3 | $2.00 / $6.00 | 4 of 6 benchmarks | 1285 | 1296 | 1328 | 1337 | ||
| 9 | Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 77.9 | $2.00 / $12.00 | 6 of 6 benchmarks | 1277 | 1294 | 1309 | 1311 | 1286 | 1274 | |
| 10 | Claude Opus 4.8 (thinking)anthropic/claude-opus-4.8:thinking | Anthropic | 77.9 | $5.00 / $25.00 | 4 of 6 benchmarks | 1284 | 1299 | 1320 | 1338 | ||
| 11 | GPT-5.5openai/gpt-5.5 | OpenAI | 77.6 | $5.00 / $30.00 | 4 of 6 benchmarks | 1286 | 1299 | 1326 | 1330 | ||
| 12 | GPT-5.5 (high)openai/gpt-5.5:high | OpenAI | 76.9 | $5.00 / $30.00 | 4 of 6 benchmarks | 1283 | 1299 | 1314 | 1339 | ||
| 13 | Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | 75.9 | $5.00 / $25.00 | 5 of 6 benchmarks | 1293 | 1310 | 1331 | 1314 | 1222 | |
| 14 | Gemini 3.6 Flashgoogle/gemini-3.6-flash | 75.1 | $1.50 / $7.50 | 3 of 6 benchmarks | 1295 | 1308 | 1313 | ||||
| 15 | Muse Sparkmeta/muse-spark | Meta | 74.4 | — | 4 of 6 benchmarks | 1294 | 1302 | 1313 | 1300 | ||
| 16 | Muse Spark 1.2 (xhigh)meta/muse-spark-1.2:xhigh | Meta | 74.4 | $1.25 / $4.25 | 3 of 6 benchmarks | 1290 | 1299 | 1316 | |||
| 17 | Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium | 72.8 | $1.50 / $9.00 | 4 of 6 benchmarks | 1284 | 1296 | 1297 | 1322 | |||
| 18 | GPT-5.4openai/gpt-5.4 | OpenAI | 72.6 | $2.50 / $15.00 | 4 of 6 benchmarks | 1280 | 1294 | 1318 | 1307 | ||
| 19 | Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 72.4 | $0.50 / $3.00 | 6 of 6 benchmarks | 1271 | 1285 | 1294 | 1307 | 1293 | 1226 | |
| 20 | Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high | Anthropic | 72.0 | $2.00 / $10.00 | 4 of 6 benchmarks | 1273 | 1291 | 1320 | 1313 | ||
| 21 | Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | 71.6 | $5.00 / $25.00 | 4 of 6 benchmarks | 1280 | 1293 | 1304 | 1322 | ||
| 22 | GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh | OpenAI | 70.7 | $5.00 / $30.00 | 4 of 6 benchmarks | 1280 | 1288 | 1308 | 1313 | ||
| 23 | GPT-5.4 (high)openai/gpt-5.4:high | OpenAI | 70.1 | $2.50 / $15.00 | 5 of 6 benchmarks | 1283 | 1299 | 1319 | 1330 | 1183 | |
| 24 | GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh | OpenAI | 69.8 | $1.00 / $6.00 | 4 of 6 benchmarks | 1270 | 1280 | 1312 | 1317 | ||
| 25 | Gemini 3.5 Flash (high)google/gemini-3.5-flash:high | 69.6 | $1.50 / $9.00 | 4 of 6 benchmarks | 1283 | 1292 | 1304 | 1294 | |||
| 26 | Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal | 68.7 | $0.50 / $3.00 | 6 of 6 benchmarks | 1259 | 1271 | 1284 | 1298 | 1278 | 1213 | |
| 27 | GPT 5.5 Instantopenai/gpt-5.5-instant | OpenAI | 67.7 | — | 4 of 6 benchmarks | 1278 | 1286 | 1315 | 1282 | ||
| 28 | GPT-5.2 Chatopenai/gpt-5.2-chat | OpenAI | 67.3 | $1.75 / $14.00 | 5 of 6 benchmarks | 1278 | 1288 | 1310 | 1293 | 1223 | |
| 29 | Muse Spark 1.1meta/muse-spark-1.1 | Meta | 67.3 | $1.25 / $4.25 | 4 of 6 benchmarks | 1282 | 1293 | 1299 | 1279 | ||
| 30 | Kimi K2.6moonshotai/kimi-k2.6 | MoonshotAI | 66.1 | $0.95 / $4.00 | 4 of 6 benchmarks | 1263 | 1278 | 1292 | 1302 | ||
| 31 | GPT-5.1 (high)openai/gpt-5.1:high | OpenAI | 64.7 | $1.25 / $10.00 | 6 of 6 benchmarks | 1250 | 1258 | 1275 | 1282 | 1240 | 1240 |
| 32 | Gemini 2.5 Progoogle/gemini-2.5-pro | 63.8 | $1.25 / $10.00 | 6 of 6 benchmarks | 1246 | 1258 | 1266 | 1272 | 1252 | 1251 | |
| 33 | Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 63.2 | $0.30 / $2.50 | 3 of 6 benchmarks | 1260 | 1282 | 1283 | ||||
| 34 | Dola Seed 2.0 Probytedance/dola-seed-2.0-pro | ByteDance | 62.0 | — | 4 of 6 benchmarks | 1258 | 1270 | 1287 | 1279 | ||
| 35 | Qwen3.7 Plusqwen/qwen3.7-plus | Qwen | 61.6 | $0.32 / $1.28 | 4 of 6 benchmarks | 1262 | 1278 | 1279 | 1276 | ||
| 36 | Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning | xAI | 61.4 | — | 5 of 6 benchmarks | 1255 | 1264 | 1273 | 1263 | 1247 | |
| 37 | Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | 61.3 | $3.00 / $15.00 | 5 of 6 benchmarks | 1275 | 1291 | 1306 | 1301 | 1159 | |
| 38 | Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking | MoonshotAI | 60.8 | $0.57 / $2.85 | 6 of 6 benchmarks | 1249 | 1264 | 1277 | 1290 | 1230 | 1196 |
| 39 | Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | Qwen | 60.2 | $0.50 / $3.60 | 5 of 6 benchmarks | 1247 | 1262 | 1276 | 1283 | 1226 | |
| 40 | Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 60.0 | $0.25 / $1.50 | 6 of 6 benchmarks | 1234 | 1247 | 1251 | 1270 | 1272 | 1226 | |
| 41 | GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh | OpenAI | 59.9 | $0.10 / $0.60 | 4 of 6 benchmarks | 1249 | 1256 | 1281 | 1285 | ||
| 42 | GPT-5.2 (high)openai/gpt-5.2:high | OpenAI | 59.7 | $1.75 / $14.00 | 6 of 6 benchmarks | 1244 | 1260 | 1274 | 1284 | 1185 | 1268 |
| 43 | Grok 4.20 Multi Agent Beta 0309x-ai/grok-4.20-multi-agent-beta-0309 | xAI | 59.6 | — | 5 of 6 benchmarks | 1252 | 1261 | 1268 | 1273 | 1231 | |
| 44 | GPT-5.4 Mini (high)openai/gpt-5.4-mini:high | OpenAI | 57.1 | $0.75 / $4.50 | 5 of 6 benchmarks | 1253 | 1266 | 1284 | 1293 | 1182 | |
| 45 | Grok 4.3x-ai/grok-4.3 | SpaceXAI | 55.6 | $1.25 / $2.50 | 5 of 6 benchmarks | 1244 | 1253 | 1268 | 1256 | 1227 | |
| 46 | Gemma 4 31Bgoogle/gemma-4-31b-it | 55.3 | $0.10 / $0.34 | 6 of 6 benchmarks | 1256 | 1270 | 1281 | 1300 | 1191 | 1165 | |
| 47 | Gemma 4 26B A4B google/gemma-4-26b-a4b-it | 54.6 | $0.12 / $0.40 | 6 of 6 benchmarks | 1240 | 1256 | 1261 | 1285 | 1153 | 1248 | |
| 48 | Chatgpt 4oopenai/chatgpt-4o | OpenAI | 54.5 | — | 6 of 6 benchmarks | 1241 | 1249 | 1264 | 1247 | 1225 | 1212 |
| 49 | MiniMax M3minimax/minimax-m3 | MiniMax | 53.7 | $0.30 / $1.20 | 4 of 6 benchmarks | 1240 | 1253 | 1263 | 1263 | ||
| 50 | GPT 4.5 Previewopenai/gpt-4.5-preview | OpenAI | 52.5 | — | 1 of 6 benchmarks | 1226 | |||||
| 51 | GPT-5 (high)openai/gpt-5:high | OpenAI | 52.4 | $1.25 / $10.00 | 6 of 6 benchmarks | 1213 | 1232 | 1252 | 1260 | 1257 | 1191 |
| 52 | GPT 5 Chatopenai/gpt-5-chat | OpenAI | 50.6 | — | 6 of 6 benchmarks | 1225 | 1245 | 1260 | 1274 | 1188 | 1198 |
| 53 | Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant | Moonshot AI | 49.0 | — | 4 of 6 benchmarks | 1237 | 1245 | 1249 | 1247 | ||
| 54 | MiMo-V2.5xiaomi/mimo-v2.5 | Xiaomi | 48.8 | $0.14 / $0.28 | 5 of 6 benchmarks | 1238 | 1253 | 1271 | 1260 | 1172 | |
| 55 | Gemini 2.5 Flashgoogle/gemini-2.5-flash | 48.4 | $0.30 / $2.50 | 6 of 6 benchmarks | 1214 | 1222 | 1233 | 1238 | 1220 | 1215 | |
| 56 | Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | Qwen | 48.1 | $0.29 / $2.40 | 4 of 6 benchmarks | 1227 | 1239 | 1246 | 1252 | ||
| 57 | GPT-5.1openai/gpt-5.1 | OpenAI | 47.9 | $1.25 / $10.00 | 6 of 6 benchmarks | 1238 | 1252 | 1264 | 1254 | 1188 | 1180 |
| 58 | Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct | Qwen | 47.7 | $0.26 / $1.04 | 6 of 6 benchmarks | 1215 | 1229 | 1239 | 1260 | 1190 | 1208 |
| 59 | o1openai/o1 | OpenAI | 47.2 | $15.00 / $60.00 | 1 of 6 benchmarks | 1193 | |||||
| 60 | o3openai/o3 | OpenAI | 46.8 | $2.00 / $8.00 | 6 of 6 benchmarks | 1217 | 1223 | 1231 | 1247 | 1234 | 1181 |
| 61 | GPT-5.2openai/gpt-5.2 | OpenAI | 46.5 | $1.75 / $14.00 | 6 of 6 benchmarks | 1229 | 1238 | 1257 | 1267 | 1192 | 1167 |
| 62 | Mimo v2 Omnixiaomi/mimo-v2-omni | Xiaomi | 46.5 | — | 4 of 6 benchmarks | 1217 | 1231 | 1254 | 1246 | ||
| 63 | GLM 5V Turboz-ai/glm-5v-turbo | Z.ai | 46.4 | $1.20 / $4.00 | 5 of 6 benchmarks | 1231 | 1244 | 1264 | 1254 | 1181 | |
| 64 | Gemini 1.5 Pro 002google/gemini-1.5-pro-002 | 44.9 | — | 1 of 6 benchmarks | 1180 | ||||||
| 65 | GPT-4.1openai/gpt-4.1 | OpenAI | 44.2 | $2.00 / $8.00 | 6 of 6 benchmarks | 1214 | 1226 | 1232 | 1249 | 1183 | 1207 |
| 66 | GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | OpenAI | 43.5 | $5.00 / $15.00 | 1 of 6 benchmarks | 1162 | |||||
| 67 | Ernie 5.0 Preview 1220baidu/ernie-5.0-preview-1220 | Baidu | 43.4 | — | 4 of 6 benchmarks | 1218 | 1230 | 1223 | 1237 | ||
| 68 | Claude Sonnet 4 (thinking 32K)anthropic/claude-sonnet-4:thinking-32k | Anthropic | 43.0 | $3.00 / $15.00 | 3 of 6 benchmarks | 1208 | 1216 | 1228 | |||
| 69 | o4 Miniopenai/o4-mini | OpenAI | 42.5 | $1.10 / $4.40 | 6 of 6 benchmarks | 1202 | 1210 | 1220 | 1251 | 1218 | 1182 |
| 70 | Qwen VL (max)qwen/qwen-vl:max | Qwen | 42.0 | — | 3 of 6 benchmarks | 1186 | 1214 | 1249 | |||
| 71 | Qwen3.5-27Bqwen/qwen3.5-27b | Qwen | 40.8 | $0.195 / $1.56 | 5 of 6 benchmarks | 1219 | 1232 | 1242 | 1242 | 1149 | |
| 72 | Mistral Large 3mistralai/mistral-large-3 | Mistral AI | 40.4 | — | 4 of 6 benchmarks | 1199 | 1219 | 1241 | 1221 | ||
| 73 | GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | 40.2 | $0.40 / $1.60 | 6 of 6 benchmarks | 1203 | 1206 | 1209 | 1222 | 1198 | 1195 |
| 74 | Gemini 1.5 Flash 002google/gemini-1.5-flash-002 | 40.2 | — | 1 of 6 benchmarks | 1141 | ||||||
| 75 | Mistral Medium 3.5mistralai/mistral-medium-3-5 | Mistral | 39.8 | $1.50 / $7.50 | 4 of 6 benchmarks | 1198 | 1215 | 1242 | 1216 | ||
| 76 | Gemini 2.0 Flash Lite Preview 02 05google/gemini-2.0-flash-lite-preview-02-05 | 39.6 | — | 1 of 6 benchmarks | 1136 | ||||||
| 77 | Qwen2.5 VL 72B Instructqwen/qwen2.5-vl-72b-instruct | Qwen | 38.5 | — | 1 of 6 benchmarks | 1122 | |||||
| 78 | Gemini 1.5 Pro 001google/gemini-1.5-pro-001 | 38.2 | — | 1 of 6 benchmarks | 1120 | ||||||
| 79 | Claude Opus 4 (thinking 16K)anthropic/claude-opus-4:thinking-16k | Anthropic | 38.1 | $15.00 / $75.00 | 4 of 6 benchmarks | 1207 | 1216 | 1203 | 1216 | ||
| 80 | Qwen2.5 VL 32B Instructqwen/qwen2.5-vl-32b-instruct | Qwen | 37.9 | — | 1 of 6 benchmarks | 1119 | |||||
| 81 | GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | OpenAI | 37.6 | $2.50 / $10.00 | 1 of 6 benchmarks | 1119 | |||||
| 82 | GPT-5 Mini (high)openai/gpt-5-mini:high | OpenAI | 37.6 | $0.25 / $2.00 | 6 of 6 benchmarks | 1183 | 1198 | 1214 | 1233 | 1207 | 1168 |
| 83 | GPT-4 Turboopenai/gpt-4-turbo | OpenAI | 37.4 | $10.00 / $30.00 | 1 of 6 benchmarks | 1112 | |||||
| 84 | Claude 3.7 Sonnet (thinking 32K)anthropic/claude-3-7-sonnet:thinking-32k | Anthropic | 37.1 | — | 4 of 6 benchmarks | 1196 | 1210 | 1216 | 1210 | ||
| 85 | GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | OpenAI | 36.8 | $0.15 / $0.60 | 1 of 6 benchmarks | 1098 | |||||
| 86 | Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking | Qwen | 36.7 | $0.40 / $4.00 | 4 of 6 benchmarks | 1190 | 1201 | 1210 | 1227 | ||
| 87 | Claude Opus 4anthropic/claude-opus-4 | Anthropic | 36.7 | $15.00 / $75.00 | 4 of 6 benchmarks | 1189 | 1197 | 1203 | 1237 | ||
| 88 | GPT-4.1 Nanoopenai/gpt-4.1-nano | OpenAI | 36.5 | $0.10 / $0.40 | 1 of 6 benchmarks | 1089 | |||||
| 89 | Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning | xAI | 36.4 | — | 5 of 6 benchmarks | 1195 | 1192 | 1215 | 1170 | 1202 | |
| 90 | Gemini 1.5 Flash 8B 001google/gemini-1.5-flash-8b-001 | 36.3 | — | 1 of 6 benchmarks | 1072 | ||||||
| 91 | Claude 3 Opusanthropic/claude-3-opus | Anthropic | 36.0 | — | 1 of 6 benchmarks | 1062 | |||||
| 92 | Gemini 1.5 Flash 001google/gemini-1.5-flash-001 | 35.7 | — | 1 of 6 benchmarks | 1060 | ||||||
| 93 | Grok 4 0709x-ai/grok-4-0709 | xAI | 35.6 | — | 6 of 6 benchmarks | 1182 | 1175 | 1191 | 1169 | 1236 | 1167 |
| 94 | Amazon Nova Pro v1.0amazon/amazon-nova-pro-v1.0 | Amazon | 35.4 | — | 1 of 6 benchmarks | 1019 | |||||
| 95 | Amazon Nova Lite v1.0amazon/amazon-nova-lite-v1.0 | Amazon | 35.1 | — | 1 of 6 benchmarks | 1019 | |||||
| 96 | Claude Sonnet 4anthropic/claude-sonnet-4 | Anthropic | 35.1 | $3.00 / $15.00 | 4 of 6 benchmarks | 1189 | 1190 | 1202 | 1222 | ||
| 97 | Claude 3 Sonnetanthropic/claude-3-sonnet | Anthropic | 34.8 | — | 1 of 6 benchmarks | 1017 | |||||
| 98 | GPT-5.4 Nano (high)openai/gpt-5.4-nano:high | OpenAI | 34.7 | $0.20 / $1.25 | 5 of 6 benchmarks | 1202 | 1215 | 1229 | 1235 | 1119 | |
| 99 | Claude 3 Haikuanthropic/claude-3-haiku | Anthropic | 34.6 | $0.25 / $1.25 | 1 of 6 benchmarks | 1001 | |||||
| 100 | Gemini 2.5 Flash Lite (thinking)google/gemini-2.5-flash-lite:thinking | 33.9 | $0.10 / $0.40 | 6 of 6 benchmarks | 1188 | 1188 | 1187 | 1190 | 1185 | 1185 | |
| 101 | Claude 3.7 Sonnetanthropic/claude-3-7-sonnet | Anthropic | 33.9 | — | 4 of 6 benchmarks | 1176 | 1186 | 1192 | 1228 | ||
| 102 | Hunyuan Vision 1.5 (thinking)tencent/hunyuan-vision-1.5:thinking | Tencent | 31.3 | — | 4 of 6 benchmarks | 1159 | 1161 | 1186 | 1229 | ||
| 103 | Gemini 2.5 Flash Lite (nothinking)google/gemini-2.5-flash-lite:nothinking | 31.2 | $0.10 / $0.40 | 4 of 6 benchmarks | 1174 | 1185 | 1183 | 1194 | |||
| 104 | Claude 3.5 Sonnetanthropic/claude-3-5-sonnet | Anthropic | 29.2 | — | 4 of 6 benchmarks | 1161 | 1177 | 1184 | 1170 | ||
| 105 | GLM 4.6Vz-ai/glm-4.6v | Z.ai | 27.8 | $0.30 / $0.90 | 4 of 6 benchmarks | 1164 | 1171 | 1186 | 1155 | ||
| 106 | Gemini 2.0 Flash 001google/gemini-2.0-flash-001 | 27.7 | — | 5 of 6 benchmarks | 1172 | 1166 | 1176 | 1190 | 1148 | ||
| 107 | GPT-5 Nano (high)openai/gpt-5-nano:high | OpenAI | 26.1 | $0.05 / $0.40 | 4 of 6 benchmarks | 1148 | 1158 | 1153 | 1187 | ||
| 108 | Hunyuan Large Visiontencent/hunyuan-large-vision | Tencent | 26.0 | — | 3 of 6 benchmarks | 1150 | 1146 | 1114 | |||
| 109 | GLM 4.5Vz-ai/glm-4.5v | Z.ai | 25.9 | $0.60 / $1.80 | 4 of 6 benchmarks | 1156 | 1158 | 1169 | 1170 | ||
| 110 | Step 3stepfun/step-3 | StepFun | 24.4 | — | 4 of 6 benchmarks | 1145 | 1151 | 1146 | 1176 | ||
| 111 | Gemma 3 27Bgoogle/gemma-3-27b-it | 24.0 | $0.08 / $0.45 | 6 of 6 benchmarks | 1160 | 1162 | 1166 | 1172 | 1175 | 1118 | |
| 112 | Step 1o Turbo 202506stepfun/step-1o-turbo-202506 | StepFun | 23.9 | — | 4 of 6 benchmarks | 1157 | 1154 | 1162 | 1138 | ||
| 113 | Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct | Meta | 23.6 | — | 4 of 6 benchmarks | 1147 | 1153 | 1152 | 1163 | ||
| 114 | Mistral Medium 2508mistralai/mistral-medium-2508 | Mistral AI | 23.0 | — | 6 of 6 benchmarks | 1159 | 1172 | 1178 | 1170 | 1132 | 1130 |
| 115 | Molmo 2 8Ballenai/molmo-2-8b | Allen Institute for AI | 22.3 | — | 3 of 6 benchmarks | 1108 | 1087 | 1108 | |||
| 116 | Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct | Meta | 21.1 | — | 4 of 6 benchmarks | 1128 | 1130 | 1148 | 1149 | ||
| 117 | Claude 3.5 Haikuanthropic/claude-3-5-haiku | Anthropic | 20.7 | — | 4 of 6 benchmarks | 1128 | 1129 | 1135 | 1150 | ||
| 118 | Mistral Medium 2505mistralai/mistral-medium-2505 | Mistral AI | 20.0 | — | 6 of 6 benchmarks | 1156 | 1159 | 1167 | 1165 | 1123 | 1092 |
| 119 | Mistral Small 2506mistralai/mistral-small-2506 | Mistral AI | 19.0 | — | 6 of 6 benchmarks | 1141 | 1142 | 1140 | 1176 | 1134 | 1055 |
| 120 | Mistral Small 3.1 24B Instruct 2503mistralai/mistral-small-3.1-24b-instruct-2503 | Mistral AI | 17.9 | — | 6 of 6 benchmarks | 1128 | 1133 | 1130 | 1159 | 1127 | 1135 |
How this ranks
Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.
A model scored on fewer than 3 of the 6 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.
Data sources
Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.
- LMArenaCC BY 4.0
Arena ratings by LMArena, from the public leaderboard dataset.