Best LLM for coding
Ranked on issue resolution, webdev preference and security review. Every row prints its coverage, and a model these benchmarks have only partly reached still ranks, on the results it does have, under a partial coverage mark.
Rumeqo runs these models inside your team rooms. See what each one costs.
| rank | model | vendor | composite | pricein / out | benchmarks | % | % | % | % | % | % | % | elo | elo | % | count |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (max)anthropic/claude-opus-5:max | Anthropic | 81.0 | $5.00 / $25.00 | 3 of 11 benchmarks | 50.0% | 1530 | 1691 | ||||||||
| 2 | Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | 79.5 | $5.00 / $25.00 | 4 of 11 benchmarks | 78.7% | 72.0% | 1547 | 1537 | |||||||
| 3 | Kimi K3 (max)moonshotai/kimi-k3:max | MoonshotAI | 78.0 | $3.00 / $15.00 | 4 of 11 benchmarks | 1543 | 1674 | 37.2% | 44.0 | |||||||
| 4 | Claude Opus 4.5anthropic/claude-opus-4.5 | Anthropic | 76.8 | $5.00 / $25.00 | 5 of 11 benchmarks | 76.7% | 52.6% | 70.7% | 1523 | 1468 | ||||||
| 5 | Claude Fable 5anthropic/claude-fable-5 | Anthropic | 76.4 | $10.00 / $50.00 | 2 of 11 benchmarks | 1554 | 1627 | |||||||||
| 6 | Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high | Anthropic | 75.8 | $5.00 / $25.00 | 3 of 11 benchmarks | 31.1% | 1552 | 1557 | ||||||||
| 7 | Claude Opus 5 (high)anthropic/claude-opus-5:high | Anthropic | 75.7 | $5.00 / $25.00 | 2 of 11 benchmarks | 1531 | 1664 | |||||||||
| 8 | Qwen3.8 Maxqwen/qwen3.8-max | Qwen | 75.6 | $2.00 / $6.00 | 2 of 11 benchmarks | 1529 | 1669 | |||||||||
| 9 | Claude Opus 4.7anthropic/claude-opus-4.7 | Anthropic | 74.7 | $5.00 / $25.00 | 2 of 11 benchmarks | 1547 | 1558 | |||||||||
| 10 | GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh | OpenAI | 74.7 | $5.00 / $30.00 | 2 of 11 benchmarks | 1526 | 1622 | |||||||||
| 11 | GPT-5.6 Luna (high)openai/gpt-5.6-luna:high | OpenAI | 73.5 | $0.10 / $0.60 | 2 of 11 benchmarks | 41.9% | 57.0 | |||||||||
| 12 | Muse Spark 1.1meta/muse-spark-1.1 | Meta | 73.1 | $1.25 / $4.25 | 2 of 11 benchmarks | 1531 | 1539 | |||||||||
| 13 | Muse Spark 1.2 (xhigh)meta/muse-spark-1.2:xhigh | Meta | 72.4 | $1.25 / $4.25 | 2 of 11 benchmarks | 1533 | 1535 | |||||||||
| 14 | Grok 4.5x-ai/grok-4.5 | SpaceXAI | 72.2 | $2.00 / $6.00 | 2 of 11 benchmarks | 1521 | 1553 | |||||||||
| 15 | Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high | Anthropic | 72.2 | $2.00 / $10.00 | 2 of 11 benchmarks | 1523 | 1541 | |||||||||
| 16 | Grok 4.6 (high)x-ai/grok-4.6:high | SpaceXAI | 71.9 | $2.00 / $6.00 | 2 of 11 benchmarks | 1512 | 1618 | |||||||||
| 17 | Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | 71.9 | $5.00 / $25.00 | 2 of 11 benchmarks | 1524 | 1539 | |||||||||
| 18 | Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 71.6 | $0.50 / $3.00 | 4 of 11 benchmarks | 75.4% | 72.7% | 1508 | 1438 | ||||||||
| 19 | Gemini 3.6 Flash (high)google/gemini-3.6-flash:high | 71.4 | $1.50 / $7.50 | 2 of 11 benchmarks | 1522 | 1537 | ||||||||||
| 20 | GLM 5.1z-ai/glm-5.1 | Z.ai | 70.0 | $1.40 / $4.40 | 3 of 11 benchmarks | 74.2% | 1515 | 1511 | ||||||||
| 21 | GPT-5.5 (high)openai/gpt-5.5:high | OpenAI | 69.8 | $5.00 / $30.00 | 5 of 11 benchmarks | 10.0% | 1520 | 1486 | 47.7% | 72.0 | ||||||
| 22 | Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | 69.8 | $3.00 / $15.00 | 5 of 11 benchmarks | 75.2% | 1528 | 1524 | 29.1% | 32.0 | ||||||
| 23 | Claude Fable 5 (high)anthropic/claude-fable-5:high | Anthropic | 69.7 | $10.00 / $50.00 | 1 of 11 benchmarks | 63.9% | ||||||||||
| 24 | Qwen3.6 Max Previewqwen/qwen3.6-max-preview | Qwen | 69.5 | $1.027 / $6.162 | 3 of 11 benchmarks | 76.7% | 1509 | 1479 | ||||||||
| 25 | GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh | OpenAI | 69.4 | $1.00 / $6.00 | 2 of 11 benchmarks | 1516 | 1523 | |||||||||
| 26 | Claude Opus 4.5 (high 32K)anthropic/claude-opus-4.5:high-32k | Anthropic | 69.3 | $5.00 / $25.00 | 2 of 11 benchmarks | 1530 | 1494 | |||||||||
| 27 | GPT 5.5 Pre Release (xhigh)openai/gpt-5.5-pre-release:xhigh | OpenAI | 69.3 | — | 1 of 11 benchmarks | 80.6% | ||||||||||
| 28 | Claude Opus 4.5 (medium)anthropic/claude-opus-4.5:medium | Anthropic | 68.6 | $5.00 / $25.00 | 1 of 11 benchmarks | 79.2% | ||||||||||
| 29 | Doubao-Seed-Codebytedance/doubao-seed-code | ByteDance | 68.2 | — | 1 of 11 benchmarks | 78.8% | ||||||||||
| 30 | Claude Fable 5 (max)anthropic/claude-fable-5:max | Anthropic | 67.8 | $10.00 / $50.00 | 1 of 11 benchmarks | 39.5% | ||||||||||
| 31 | Grok 4.5 (high)x-ai/grok-4.5:high | SpaceXAI | 67.7 | $2.00 / $6.00 | 2 of 11 benchmarks | 38.4% | 41.0 | |||||||||
| 32 | Muse Sparkmeta/muse-spark | Meta | 67.7 | — | 1 of 11 benchmarks | 1526 | ||||||||||
| 33 | GLM 5.2 (max)z-ai/glm-5.2:max | Z.ai | 67.1 | $0.63 / $1.98 | 4 of 11 benchmarks | 78.7% | 9.5% | 1506 | 1587 | |||||||
| 34 | DeepSeek V4 Pro (max)deepseek/deepseek-v4-pro:max | DeepSeek | 67.1 | $1.168 / $2.336 | 1 of 11 benchmarks | 77.6% | ||||||||||
| 35 | Hy3tencent/hy3 | Tencent | 67.1 | $0.132 / $0.528 | 2 of 11 benchmarks | 1503 | 1523 | |||||||||
| 36 | Claude Opus 4.7 (max)anthropic/claude-opus-4.7:max | Anthropic | 67.0 | $5.00 / $25.00 | 2 of 11 benchmarks | 83.5% | 19.1% | |||||||||
| 37 | DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high | DeepSeek | 66.9 | $0.14 / $0.28 | 2 of 11 benchmarks | 1480 | 1582 | |||||||||
| 38 | MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | Xiaomi | 66.9 | $0.435 / $0.87 | 2 of 11 benchmarks | 1520 | 1474 | |||||||||
| 39 | GPT-5.5 (xhigh)openai/gpt-5.5:xhigh | OpenAI | 66.5 | $5.00 / $30.00 | 2 of 11 benchmarks | 34.3% | 1509 | |||||||||
| 40 | Qwen3.7 Maxqwen/qwen3.7-max | Qwen | 66.2 | $1.475 / $4.425 | 4 of 11 benchmarks | 77.3% | 9.5% | 1525 | 1517 | |||||||
| 41 | Claude Opus 4.5 (high)anthropic/claude-opus-4.5:high | Anthropic | 66.0 | $5.00 / $25.00 | 1 of 11 benchmarks | 76.8% | ||||||||||
| 42 | GPT-5.6 Sol (max)openai/gpt-5.6-sol:max | OpenAI | 65.8 | $5.00 / $30.00 | 1 of 11 benchmarks | 39.0% | ||||||||||
| 43 | GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh | OpenAI | 65.8 | $0.10 / $0.60 | 2 of 11 benchmarks | 1499 | 1518 | |||||||||
| 44 | GPT-5.4 (high)openai/gpt-5.4:high | OpenAI | 65.7 | $2.50 / $15.00 | 4 of 11 benchmarks | 76.9% | 15.6% | 1521 | 1463 | |||||||
| 45 | GPT-5.2 Chatopenai/gpt-5.2-chat | OpenAI | 65.6 | $1.75 / $14.00 | 1 of 11 benchmarks | 1515 | ||||||||||
| 46 | Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium | 65.3 | $1.50 / $9.00 | 2 of 11 benchmarks | 1507 | 1488 | ||||||||||
| 47 | Ernie 5.1baidu/ernie-5.1 | Baidu | 65.3 | — | 1 of 11 benchmarks | 1514 | ||||||||||
| 48 | GPT 5.5 Instantopenai/gpt-5.5-instant | OpenAI | 65.3 | — | 1 of 11 benchmarks | 1514 | ||||||||||
| 49 | Qwen3.5 Max Previewqwen/qwen3.5-max-preview | Qwen | 64.9 | — | 1 of 11 benchmarks | 1513 | ||||||||||
| 50 | Dola Seed 2.0 Probytedance/dola-seed-2.0-pro | ByteDance | 64.7 | — | 1 of 11 benchmarks | 1513 | ||||||||||
| 51 | Claude Opus 4.1 (thinking 16K)anthropic/claude-opus-4.1:thinking-16k | Anthropic | 64.5 | $15.00 / $75.00 | 1 of 11 benchmarks | 1512 | ||||||||||
| 52 | Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high | 64.2 | $0.50 / $3.00 | 1 of 11 benchmarks | 75.8% | |||||||||||
| 53 | MiniMax M2.5 (high)minimax/minimax-m2.5:high | MiniMax | 64.2 | $0.22 / $0.90 | 1 of 11 benchmarks | 75.8% | ||||||||||
| 54 | Gemini 3 Progoogle/gemini-3-pro | 64.2 | — | 3 of 11 benchmarks | 68.7% | 1518 | 1438 | |||||||||
| 55 | MiniMax M3minimax/minimax-m3 | MiniMax | 64.1 | $0.30 / $1.20 | 2 of 11 benchmarks | 1497 | 1491 | |||||||||
| 56 | GPT-5.5openai/gpt-5.5 | OpenAI | 63.8 | $5.00 / $30.00 | 2 of 11 benchmarks | 1510 | 1458 | |||||||||
| 57 | Grok 4.20 Multi Agent Beta 0309x-ai/grok-4.20-multi-agent-beta-0309 | xAI | 63.7 | — | 1 of 11 benchmarks | 1508 | ||||||||||
| 58 | Gemini 3.1 Pro Preview Custom Toolsgoogle/gemini-3.1-pro-preview-customtools | 63.7 | $2.00 / $12.00 | 1 of 11 benchmarks | 75.6% | |||||||||||
| 59 | GLM 5z-ai/glm-5 | Z.ai | 63.4 | $0.95 / $2.55 | 4 of 11 benchmarks | 72.1% | 69.7% | 1497 | 1436 | |||||||
| 60 | Qwen3.7 Plusqwen/qwen3.7-plus | Qwen | 63.3 | $0.32 / $1.28 | 1 of 11 benchmarks | 1506 | ||||||||||
| 61 | Claude Sonnet 4anthropic/claude-sonnet-4 | Anthropic | 63.2 | $3.00 / $15.00 | 4 of 11 benchmarks | 57.0% | 58.3% | 35.6% | 1449 | |||||||
| 62 | Seed 2.1 Pro Previewbytedance/seed-2.1-pro-preview | ByteDance | 62.8 | — | 1 of 11 benchmarks | 1522 | ||||||||||
| 63 | GPT-5.3-Codex (high)openai/gpt-5.3-codex:high | OpenAI | 62.5 | $1.75 / $14.00 | 1 of 11 benchmarks | 74.8% | ||||||||||
| 64 | Gemini 3.5 Flash (high)google/gemini-3.5-flash:high | 62.4 | $1.50 / $9.00 | 4 of 11 benchmarks | 79.3% | 4.8% | 1509 | 1506 | ||||||||
| 65 | Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 62.4 | $0.30 / $2.50 | 2 of 11 benchmarks | 1503 | 1449 | ||||||||||
| 66 | Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 62.3 | $2.00 / $12.00 | 3 of 11 benchmarks | 14.3% | 1521 | 1447 | |||||||||
| 67 | Claude Opus 4.6 (high)anthropic/claude-opus-4.6:high | Anthropic | 62.3 | $5.00 / $25.00 | 4 of 11 benchmarks | 1552 | 1545 | 26.7% | 24.0 | |||||||
| 68 | GPT-5openai/gpt-5 | OpenAI | 62.2 | $1.25 / $10.00 | 1 of 11 benchmarks | 74.4% | ||||||||||
| 69 | Claude Opus 4 (thinking 16K)anthropic/claude-opus-4:thinking-16k | Anthropic | 62.1 | $15.00 / $75.00 | 1 of 11 benchmarks | 1499 | ||||||||||
| 70 | GPT-5.5 (low)openai/gpt-5.5:low | OpenAI | 61.9 | $5.00 / $30.00 | 2 of 11 benchmarks | 32.6% | 38.0 | |||||||||
| 71 | Claude Opus 4.8 (max)anthropic/claude-opus-4.8:max | Anthropic | 61.9 | $5.00 / $25.00 | 1 of 11 benchmarks | 28.6% | ||||||||||
| 72 | DeepSeek V4 Prodeepseek/deepseek-v4-pro | DeepSeek | 61.7 | $1.168 / $2.336 | 2 of 11 benchmarks | 1502 | 1445 | |||||||||
| 73 | GPT-5.3 Chatopenai/gpt-5.3-chat | OpenAI | 61.1 | — | 1 of 11 benchmarks | 1496 | ||||||||||
| 74 | DeepSeek V4 Pro (high)deepseek/deepseek-v4-pro:high | DeepSeek | 61.1 | $1.168 / $2.336 | 2 of 11 benchmarks | 1489 | 1464 | |||||||||
| 75 | Grok 4.1x-ai/grok-4.1 | xAI | 60.7 | — | 1 of 11 benchmarks | 1492 | ||||||||||
| 76 | o3openai/o3 | OpenAI | 60.6 | $2.00 / $8.00 | 3 of 11 benchmarks | 58.4% | 36.0% | 1460 | ||||||||
| 77 | Claude Sonnet 4.5 (high 32K)anthropic/claude-sonnet-4.5:high-32k | Anthropic | 60.5 | $3.00 / $15.00 | 2 of 11 benchmarks | 1519 | 1392 | |||||||||
| 78 | Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking | MoonshotAI | 60.5 | $0.57 / $2.85 | 2 of 11 benchmarks | 1502 | 1436 | |||||||||
| 79 | Mimo v2 Proxiaomi/mimo-v2-pro | Xiaomi | 60.2 | — | 2 of 11 benchmarks | 1503 | 1434 | |||||||||
| 80 | Kimi K2.6moonshotai/kimi-k2.6 | MoonshotAI | 60.2 | $0.95 / $4.00 | 4 of 11 benchmarks | 76.7% | 2.4% | 1514 | 1509 | |||||||
| 81 | GPT-5.4 (xhigh)openai/gpt-5.4:xhigh | OpenAI | 59.9 | $2.50 / $15.00 | 1 of 11 benchmarks | 25.4% | ||||||||||
| 82 | Gemini 3 Pro Previewgoogle/gemini-3-pro-preview | 59.9 | — | 1 of 11 benchmarks | 72.9% | |||||||||||
| 83 | Ernie 5.0 0110baidu/ernie-5.0-0110 | Baidu | 59.7 | — | 1 of 11 benchmarks | 1490 | ||||||||||
| 84 | Claude 3.7 Sonnetanthropic/claude-3-7-sonnet | Anthropic | 59.6 | — | 5 of 11 benchmarks | 61.0% | 51.7% | 33.8% | 31.3% | 1430 | ||||||
| 85 | MiMo-V2.5xiaomi/mimo-v2.5 | Xiaomi | 59.6 | $0.14 / $0.28 | 2 of 11 benchmarks | 1491 | 1438 | |||||||||
| 86 | GLM 5 (high)z-ai/glm-5:high | Z.ai | 59.3 | $0.95 / $2.55 | 1 of 11 benchmarks | 72.8% | ||||||||||
| 87 | Mimo v2 Omnixiaomi/mimo-v2-omni | Xiaomi | 59.3 | — | 1 of 11 benchmarks | 1486 | ||||||||||
| 88 | Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high | Anthropic | 59.3 | $5.00 / $25.00 | 4 of 11 benchmarks | 1533 | 1564 | 24.4% | 24.0 | |||||||
| 89 | GPT-5.4openai/gpt-5.4 | OpenAI | 59.3 | $2.50 / $15.00 | 2 of 11 benchmarks | 1514 | 1390 | |||||||||
| 90 | Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant | Moonshot AI | 59.2 | — | 2 of 11 benchmarks | 1505 | 1405 | |||||||||
| 91 | DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | DeepSeek | 58.9 | $0.14 / $0.28 | 1 of 11 benchmarks | 1483 | ||||||||||
| 92 | GPT-5.1 (high)openai/gpt-5.1:high | OpenAI | 58.8 | $1.25 / $10.00 | 2 of 11 benchmarks | 68.0% | 1491 | |||||||||
| 93 | Kimi K2.7 Codemoonshotai/kimi-k2.7-code | MoonshotAI | 58.8 | $0.67 / $3.40 | 1 of 11 benchmarks | 1473 | ||||||||||
| 94 | Kimi K2.5moonshotai/kimi-k2.5 | MoonshotAI | 58.4 | $0.57 / $2.85 | 2 of 11 benchmarks | 73.8% | 67.3% | |||||||||
| 95 | Claude Sonnet 4.5 (high)anthropic/claude-sonnet-4.5:high | Anthropic | 58.0 | $3.00 / $15.00 | 1 of 11 benchmarks | 71.4% | ||||||||||
| 96 | GPT-5.2 (xhigh)openai/gpt-5.2:xhigh | OpenAI | 58.0 | $1.75 / $14.00 | 1 of 11 benchmarks | 23.0% | ||||||||||
| 97 | Inklingthinkingmachines/inkling | Thinking Machines | 57.9 | $0.95 / $4.05 | 2 of 11 benchmarks | 1494 | 1405 | |||||||||
| 98 | Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | NVIDIA | 57.8 | $0.60 / $3.60 | 1 of 11 benchmarks | 1475 | ||||||||||
| 99 | GLM 4.7z-ai/glm-4.7 | Z.ai | 57.7 | $0.40 / $1.75 | 2 of 11 benchmarks | 1485 | 1434 | |||||||||
| 100 | Kimi K2 0905moonshotai/kimi-k2-0905 | MoonshotAI | 57.6 | $0.60 / $2.50 | 2 of 11 benchmarks | 71.2% | 1468 | |||||||||
| 101 | DeepSeek V3.2 Exp (thinking)deepseek/deepseek-v3.2-exp:thinking | DeepSeek | 57.5 | $0.27 / $0.41 | 1 of 11 benchmarks | 1475 | ||||||||||
| 102 | GPT-5.2 (high)openai/gpt-5.2:high | OpenAI | 57.5 | $1.75 / $14.00 | 3 of 11 benchmarks | 73.8% | 66.7% | 1490 | ||||||||
| 103 | Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning | xAI | 57.3 | — | 2 of 11 benchmarks | 1511 | 1374 | |||||||||
| 104 | LongCat Flash Chatmeituan/longcat-flash-chat | Meituan | 57.3 | — | 1 of 11 benchmarks | 1474 | ||||||||||
| 105 | Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | Qwen | 57.2 | $0.50 / $3.60 | 2 of 11 benchmarks | 1491 | 1400 | |||||||||
| 106 | Claude Sonnet 4 (thinking 32K)anthropic/claude-sonnet-4:thinking-32k | Anthropic | 57.1 | $3.00 / $15.00 | 1 of 11 benchmarks | 1473 | ||||||||||
| 107 | GPT-5.4 Mini (high)openai/gpt-5.4-mini:high | OpenAI | 57.1 | $0.75 / $4.50 | 2 of 11 benchmarks | 1497 | 1397 | |||||||||
| 108 | Qwen3.6 Plusqwen/qwen3.6-plus | Qwen | 57.0 | $0.325 / $1.95 | 3 of 11 benchmarks | 57.9% | 1495 | 1459 | ||||||||
| 109 | Qwen3 Maxqwen/qwen3-max | Qwen | 57.0 | $0.78 / $3.90 | 1 of 11 benchmarks | 1473 | ||||||||||
| 110 | Kimi K2.5 (high)moonshotai/kimi-k2.5:high | MoonshotAI | 56.9 | $0.57 / $2.85 | 1 of 11 benchmarks | 70.8% | ||||||||||
| 111 | Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | Qwen | 56.8 | $0.09 / $0.55 | 1 of 11 benchmarks | 1472 | ||||||||||
| 112 | GPT-5 (medium)openai/gpt-5:medium | OpenAI | 56.7 | $1.25 / $10.00 | 2 of 11 benchmarks | 71.5% | 1419 | |||||||||
| 113 | Ernie 5.0 Preview 1203baidu/ernie-5.0-preview-1203 | Baidu | 56.7 | — | 1 of 11 benchmarks | 1472 | ||||||||||
| 114 | GLM 5V Turboz-ai/glm-5v-turbo | Z.ai | 56.6 | $1.20 / $4.00 | 2 of 11 benchmarks | 1490 | 1400 | |||||||||
| 115 | GPT-5.2openai/gpt-5.2 | OpenAI | 56.6 | $1.75 / $14.00 | 3 of 11 benchmarks | 69.0% | 1482 | 1418 | ||||||||
| 116 | Claude Opus 4anthropic/claude-opus-4 | Anthropic | 56.5 | $15.00 / $75.00 | 2 of 11 benchmarks | 70.7% | 1464 | |||||||||
| 117 | GPT-5.6 Sol (high)openai/gpt-5.6-sol:high | OpenAI | 56.4 | $5.00 / $30.00 | 1 of 11 benchmarks | 20.0% | ||||||||||
| 118 | Chatgpt 4oopenai/chatgpt-4o | OpenAI | 56.3 | — | 1 of 11 benchmarks | 1468 | ||||||||||
| 119 | DeepSeek V3.2 (high)deepseek/deepseek-v3.2:high | DeepSeek | 56.1 | $0.269 / $0.40 | 1 of 11 benchmarks | 70.0% | ||||||||||
| 120 | GPT-5.4 (medium)openai/gpt-5.4:medium | OpenAI | 56.1 | $2.50 / $15.00 | 1 of 11 benchmarks | 1442 | ||||||||||
| 121 | GPT-5 (high)openai/gpt-5:high | OpenAI | 56.0 | $1.25 / $10.00 | 3 of 11 benchmarks | 73.5% | 12.7% | 1469 | ||||||||
| 122 | Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct | Qwen | 55.6 | $0.26 / $1.04 | 1 of 11 benchmarks | 1465 | ||||||||||
| 123 | Gemini 3 Pro Preview (high)google/gemini-3-pro-preview:high | 55.5 | — | 1 of 11 benchmarks | 69.6% | |||||||||||
| 124 | R1 0528deepseek/deepseek-r1-0528 | DeepSeek | 55.5 | $0.50 / $2.15 | 1 of 11 benchmarks | 1464 | ||||||||||
| 125 | DeepSeek V3.1 Terminus (thinking)deepseek/deepseek-v3.1-terminus:thinking | DeepSeek | 55.2 | $0.27 / $0.95 | 1 of 11 benchmarks | 1463 | ||||||||||
| 126 | GPT 5 Chatopenai/gpt-5-chat | OpenAI | 55.1 | — | 1 of 11 benchmarks | 1462 | ||||||||||
| 127 | MiniMax M2.7minimax/minimax-m2.7 | MiniMax | 55.0 | $0.30 / $1.20 | 2 of 11 benchmarks | 1479 | 1398 | |||||||||
| 128 | Gemma 4 31Bgoogle/gemma-4-31b-it | 55.0 | $0.10 / $0.34 | 2 of 11 benchmarks | 1499 | 1365 | ||||||||||
| 129 | GPT-5.4 Nano (high)openai/gpt-5.4-nano:high | OpenAI | 54.5 | $0.20 / $1.25 | 1 of 11 benchmarks | 1460 | ||||||||||
| 130 | Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal | 54.4 | $0.50 / $3.00 | 2 of 11 benchmarks | 1491 | 1383 | ||||||||||
| 131 | Claude Haiku 4.5 (high)anthropic/claude-haiku-4.5:high | Anthropic | 54.2 | $1.00 / $5.00 | 1 of 11 benchmarks | 66.6% | ||||||||||
| 132 | GPT 4.5 Previewopenai/gpt-4.5-preview | OpenAI | 54.1 | — | 1 of 11 benchmarks | 1459 | ||||||||||
| 133 | Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | 54.0 | $3.00 / $15.00 | 6 of 11 benchmarks | 71.3% | 44.3% | 67.0% | 2.4% | 1513 | 1386 | |||||
| 134 | GPT-5.1-Codex (medium)openai/gpt-5.1-codex:medium | OpenAI | 53.6 | $1.25 / $10.00 | 1 of 11 benchmarks | 66.0% | ||||||||||
| 135 | DeepSeek V3.1 (thinking)deepseek/deepseek-chat-v3.1:thinking | DeepSeek | 53.5 | $0.25 / $0.95 | 1 of 11 benchmarks | 1457 | ||||||||||
| 136 | Kimi K2 0711moonshotai/kimi-k2 | MoonshotAI | 53.5 | $0.57 / $2.30 | 2 of 11 benchmarks | 65.4% | 1461 | |||||||||
| 137 | Mistral Medium 2508mistralai/mistral-medium-2508 | Mistral AI | 53.3 | — | 1 of 11 benchmarks | 1455 | ||||||||||
| 138 | DeepSeek V4 Pro (xhigh)deepseek/deepseek-v4-pro:xhigh | DeepSeek | 53.3 | $1.168 / $2.336 | 2 of 11 benchmarks | 26.7% | 30.0 | |||||||||
| 139 | Claude Opus 4.1anthropic/claude-opus-4.1 | Anthropic | 53.2 | $15.00 / $75.00 | 4 of 11 benchmarks | 73.3% | 7.9% | 1505 | 1389 | |||||||
| 140 | Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking | Qwen | 53.1 | $0.40 / $4.00 | 1 of 11 benchmarks | 1455 | ||||||||||
| 141 | Claude Opus 4.5 (128K)anthropic/claude-opus-4.5:128k | Anthropic | 53.1 | $5.00 / $25.00 | 1 of 11 benchmarks | 14.3% | ||||||||||
| 142 | GPT-5.3-Codexopenai/gpt-5.3-codex | OpenAI | 53.1 | $1.75 / $14.00 | 1 of 11 benchmarks | 1409 | ||||||||||
| 143 | Claude 3.7 Sonnet (thinking 32K)anthropic/claude-3-7-sonnet:thinking-32k | Anthropic | 52.9 | — | 1 of 11 benchmarks | 1452 | ||||||||||
| 144 | Step 3.5 Flashstepfun/step-3.5-flash | StepFun | 52.7 | $0.10 / $0.30 | 1 of 11 benchmarks | 1451 | ||||||||||
| 145 | GPT-5 Mini (medium)openai/gpt-5-mini:medium | OpenAI | 52.7 | $0.25 / $2.00 | 1 of 11 benchmarks | 64.7% | ||||||||||
| 146 | Claude 3.5 Haikuanthropic/claude-3-5-haiku | Anthropic | 52.4 | — | 2 of 11 benchmarks | 41.7% | 1385 | |||||||||
| 147 | Gemma 4 26B A4B google/gemma-4-26b-a4b-it | 52.3 | $0.12 / $0.40 | 2 of 11 benchmarks | 1481 | 1362 | ||||||||||
| 148 | Claude 3.5 Sonnetanthropic/claude-3-5-sonnet | Anthropic | 52.3 | — | 5 of 11 benchmarks | 62.8% | 51.3% | 24.9% | 25.3% | 1435 | ||||||
| 149 | DeepSeek V3.1deepseek/deepseek-chat-v3.1 | DeepSeek | 52.2 | $0.25 / $0.95 | 1 of 11 benchmarks | 1448 | ||||||||||
| 150 | Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | MoonshotAI | 51.9 | $0.60 / $2.50 | 1 of 11 benchmarks | 63.4% | ||||||||||
| 151 | Muse Glimmer 30Bmeta/muse-glimmer-30b | Meta | 51.9 | $0.35 / $1.50 | 2 of 11 benchmarks | 1481 | 1359 | |||||||||
| 152 | Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct | Qwen | 51.9 | $0.10 / $1.10 | 1 of 11 benchmarks | 1446 | ||||||||||
| 153 | Qwen3 235B A22B (nothinking)qwen/qwen3-235b-a22b:nothinking | Qwen | 51.8 | $0.455 / $1.82 | 1 of 11 benchmarks | 1446 | ||||||||||
| 154 | Grok 4.3x-ai/grok-4.3 | SpaceXAI | 51.6 | $1.25 / $2.50 | 2 of 11 benchmarks | 1488 | 1355 | |||||||||
| 155 | R1deepseek/deepseek-r1 | DeepSeek | 51.6 | $0.70 / $2.50 | 1 of 11 benchmarks | 1445 | ||||||||||
| 156 | Grok 3 Betax-ai/grok-3-beta | xAI | 51.3 | — | 1 of 11 benchmarks | 1443 | ||||||||||
| 157 | Trinity Large Previewarcee-ai/trinity-large-preview | Arcee AI | 51.2 | — | 1 of 11 benchmarks | 1443 | ||||||||||
| 158 | o3 (medium)openai/o3:medium | OpenAI | 51.2 | $2.00 / $8.00 | 1 of 11 benchmarks | 62.3% | ||||||||||
| 159 | Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | Qwen | 51.1 | $0.23 / $2.30 | 1 of 11 benchmarks | 1442 | ||||||||||
| 160 | GPT-5.1 (medium)openai/gpt-5.1:medium | OpenAI | 50.9 | $1.25 / $10.00 | 2 of 11 benchmarks | 66.0% | 1391 | |||||||||
| 161 | Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | Qwen | 50.8 | $0.0482 / $0.1931 | 1 of 11 benchmarks | 1440 | ||||||||||
| 162 | DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | DeepSeek | 50.7 | $0.27 / $0.95 | 1 of 11 benchmarks | 1439 | ||||||||||
| 163 | Hunyuan Vision 1.5 (thinking)tencent/hunyuan-vision-1.5:thinking | Tencent | 50.5 | — | 1 of 11 benchmarks | 1438 | ||||||||||
| 164 | MiniMax M2.5minimax/minimax-m2.5 | MiniMax | 50.2 | $0.22 / $0.90 | 3 of 11 benchmarks | 68.3% | 1444 | 1384 | ||||||||
| 165 | Grok 4 0709x-ai/grok-4-0709 | xAI | 49.8 | — | 1 of 11 benchmarks | 1435 | ||||||||||
| 166 | o3 Mini Highopenai/o3-mini-high | OpenAI | 49.7 | $1.10 / $4.40 | 1 of 11 benchmarks | 1435 | ||||||||||
| 167 | GPT-5.1openai/gpt-5.1 | OpenAI | 49.6 | $1.25 / $10.00 | 2 of 11 benchmarks | 1474 | 1341 | |||||||||
| 168 | Mistral Medium 2505mistralai/mistral-medium-2505 | Mistral AI | 49.4 | — | 1 of 11 benchmarks | 1433 | ||||||||||
| 169 | Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo | Moonshot AI | 49.4 | — | 2 of 11 benchmarks | 1486 | 1323 | |||||||||
| 170 | DeepSeek V3deepseek/deepseek-chat | DeepSeek | 49.3 | $0.2574 / $1.0287 | 2 of 11 benchmarks | 36.7% | 1388 | |||||||||
| 171 | DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking | DeepSeek | 49.3 | $0.269 / $0.40 | 3 of 11 benchmarks | 60.0% | 1475 | 1361 | ||||||||
| 172 | Qwen3 235B A22Bqwen/qwen3-235b-a22b | Qwen | 49.2 | $0.455 / $1.82 | 1 of 11 benchmarks | 1433 | ||||||||||
| 173 | Claude Opus 4.6 (max)anthropic/claude-opus-4.6:max | Anthropic | 49.1 | $5.00 / $25.00 | 1 of 11 benchmarks | 12.7% | ||||||||||
| 174 | o4 Miniopenai/o4-mini | OpenAI | 48.9 | $1.10 / $4.40 | 3 of 11 benchmarks | 45.0% | 33.9% | 1433 | ||||||||
| 175 | o1openai/o1 | OpenAI | 48.9 | $15.00 / $60.00 | 2 of 11 benchmarks | 64.6% | 1433 | |||||||||
| 176 | Ernie 5.0 Preview 1022baidu/ernie-5.0-preview-1022 | Baidu | 48.9 | — | 1 of 11 benchmarks | 1432 | ||||||||||
| 177 | GPT-5 Mini (high)openai/gpt-5-mini:high | OpenAI | 48.6 | $0.25 / $2.00 | 1 of 11 benchmarks | 1431 | ||||||||||
| 178 | Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | Qwen | 48.4 | $0.29 / $2.40 | 2 of 11 benchmarks | 1459 | 1358 | |||||||||
| 179 | Hunyuan Hy3 Previewtencent/hunyuan-hy3-preview | Tencent | 48.4 | — | 2 of 11 benchmarks | 1461 | 1356 | |||||||||
| 180 | DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | DeepSeek | 48.3 | $0.27 / $1.12 | 1 of 11 benchmarks | 1429 | ||||||||||
| 181 | Solar Pro 4upstage/solar-pro4 | Upstage | 48.3 | $0.03 / $0.12 | 2 of 11 benchmarks | 1450 | 1372 | |||||||||
| 182 | MiniMax M2.1minimax/minimax-m2.1 | MiniMax | 48.2 | $0.30 / $1.20 | 2 of 11 benchmarks | 1440 | 1387 | |||||||||
| 183 | GLM 4.5 Airz-ai/glm-4.5-air | Z.ai | 48.2 | $0.13 / $0.85 | 1 of 11 benchmarks | 1426 | ||||||||||
| 184 | Devstral Small 2512mistralai/devstral-small-2512 | Mistral AI | 48.1 | — | 1 of 11 benchmarks | 56.4% | ||||||||||
| 185 | GLM 4.7 Flashz-ai/glm-4.7-flash | Z.ai | 48.1 | $0.06 / $0.40 | 1 of 11 benchmarks | 1424 | ||||||||||
| 186 | Grok 4.1 (thinking)x-ai/grok-4.1:thinking | xAI | 47.8 | — | 2 of 11 benchmarks | 1499 | 1210 | |||||||||
| 187 | Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning-30b-a3b | NVIDIA | 47.8 | — | 1 of 11 benchmarks | 1422 | ||||||||||
| 188 | GLM 4.5z-ai/glm-4.5 | Z.ai | 47.7 | $0.60 / $2.20 | 2 of 11 benchmarks | 54.2% | 1455 | |||||||||
| 189 | Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | Qwen | 47.6 | $0.15 / $1.20 | 1 of 11 benchmarks | 1421 | ||||||||||
| 190 | GLM 4.6Vz-ai/glm-4.6v | Z.ai | 47.5 | $0.30 / $0.90 | 1 of 11 benchmarks | 1417 | ||||||||||
| 191 | Claude Sonnet 5anthropic/claude-sonnet-5 | Anthropic | 47.5 | $2.00 / $10.00 | 2 of 11 benchmarks | 25.6% | 27.0 | |||||||||
| 192 | MiniMax M1minimax/minimax-m1 | MiniMax | 47.2 | $0.55 / $2.20 | 1 of 11 benchmarks | 1416 | ||||||||||
| 193 | Mistral Medium 3.5mistralai/mistral-medium-3-5 | Mistral | 47.1 | $1.50 / $7.50 | 2 of 11 benchmarks | 1479 | 1266 | |||||||||
| 194 | Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct | Qwen | 47.0 | — | 3 of 11 benchmarks | 69.6% | 1457 | 1273 | ||||||||
| 195 | Mistral Small 2506mistralai/mistral-small-2506 | Mistral AI | 47.0 | — | 1 of 11 benchmarks | 1412 | ||||||||||
| 196 | Qwen3.5-27Bqwen/qwen3.5-27b | Qwen | 46.8 | $0.195 / $1.56 | 2 of 11 benchmarks | 1450 | 1357 | |||||||||
| 197 | Ling Flash 2.0inclusionai/ling-flash-2.0 | inclusionAI | 46.8 | — | 1 of 11 benchmarks | 1411 | ||||||||||
| 198 | INTELLECT-3prime-intellect/intellect-3 | Prime Intellect | 46.7 | — | 1 of 11 benchmarks | 1409 | ||||||||||
| 199 | Step 3stepfun/step-3 | StepFun | 46.6 | — | 1 of 11 benchmarks | 1408 | ||||||||||
| 200 | Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | NVIDIA | 46.4 | $0.085 / $0.40 | 1 of 11 benchmarks | 1408 | ||||||||||
| 201 | GPT-4.1openai/gpt-4.1 | OpenAI | 46.3 | $2.00 / $8.00 | 3 of 11 benchmarks | 48.5% | 31.1% | 1456 | ||||||||
| 202 | Qwen3 32Bqwen/qwen3-32b | Qwen | 46.3 | $0.08 / $0.28 | 1 of 11 benchmarks | 1407 | ||||||||||
| 203 | Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | Qwen | 46.3 | $0.07 / $0.28 | 1 of 11 benchmarks | 51.6% | ||||||||||
| 204 | GLM 4.5Vz-ai/glm-4.5v | Z.ai | 46.1 | $0.60 / $1.80 | 1 of 11 benchmarks | 1405 | ||||||||||
| 205 | Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | NVIDIA | 46.0 | — | 1 of 11 benchmarks | 1404 | ||||||||||
| 206 | Qwen 2.5 (max)qwen/qwen-2.5:max | Qwen | 45.9 | — | 1 of 11 benchmarks | 1403 | ||||||||||
| 207 | Hunyuan T1tencent/hunyuan-t1 | Tencent | 45.7 | — | 1 of 11 benchmarks | 1399 | ||||||||||
| 208 | DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | DeepSeek | 45.7 | $0.27 / $0.41 | 2 of 11 benchmarks | 1465 | 1272 | |||||||||
| 209 | Gemini 2.5 Flash Lite (nothinking)google/gemini-2.5-flash-lite:nothinking | 45.6 | $0.10 / $0.40 | 1 of 11 benchmarks | 1397 | |||||||||||
| 210 | GPT-5.2-Codexopenai/gpt-5.2-codex | OpenAI | 45.5 | $1.75 / $14.00 | 3 of 11 benchmarks | 72.8% | 66.3% | 1338 | ||||||||
| 211 | Devstral Small 2505mistralai/devstral-small-2505 | Mistral AI | 45.5 | — | 1 of 11 benchmarks | 46.8% | ||||||||||
| 212 | Laguna M.1poolside/laguna-m.1 | Poolside | 45.5 | — | 1 of 11 benchmarks | 1348 | ||||||||||
| 213 | Nova 2 Liteamazon/nova-2-lite-v1 | Amazon | 45.5 | $0.30 / $2.50 | 1 of 11 benchmarks | 1395 | ||||||||||
| 214 | Hunyuan TurboStencent/hunyuan-turbos | Tencent | 45.3 | — | 1 of 11 benchmarks | 1394 | ||||||||||
| 215 | Llama 3.1 Nemotron Ultra 253B v1nvidia/llama-3.1-nemotron-ultra-253b-v1 | NVIDIA | 45.0 | — | 1 of 11 benchmarks | 1391 | ||||||||||
| 216 | Ring Flash 2.0inclusionai/ring-flash-2.0 | inclusionAI | 44.9 | — | 1 of 11 benchmarks | 1390 | ||||||||||
| 217 | Kimi K2 Instructmoonshotai/kimi-k2-instruct | Moonshot AI | 44.7 | — | 1 of 11 benchmarks | 43.8% | ||||||||||
| 218 | Mimo v2 Flashxiaomi/mimo-v2-flash | Xiaomi | 44.7 | — | 2 of 11 benchmarks | 1446 | 1330 | |||||||||
| 219 | Grok 3 Mini Beta (high)x-ai/grok-3-mini-beta:high | xAI | 44.6 | — | 1 of 11 benchmarks | 1390 | ||||||||||
| 220 | o3 Miniopenai/o3-mini | OpenAI | 44.5 | $1.10 / $4.40 | 3 of 11 benchmarks | 42.4% | 32.3% | 1416 | ||||||||
| 221 | Command A (03-2025)cohere/command-a-03-2025 | Cohere | 44.5 | — | 1 of 11 benchmarks | 1390 | ||||||||||
| 222 | GPT-5.1-Codexopenai/gpt-5.1-codex | OpenAI | 44.3 | $1.25 / $10.00 | 1 of 11 benchmarks | 1336 | ||||||||||
| 223 | Magistral Medium 2506mistralai/magistral-medium-2506 | Mistral AI | 44.2 | — | 1 of 11 benchmarks | 1387 | ||||||||||
| 224 | Nova Premier 1.0amazon/nova-premier-v1 | Amazon | 44.2 | $2.50 / $12.50 | 1 of 11 benchmarks | 42.4% | ||||||||||
| 225 | O1 Miniopenai/o1-mini | OpenAI | 44.1 | — | 1 of 11 benchmarks | 1387 | ||||||||||
| 226 | GLM 4.6z-ai/glm-4.6 | Z.ai | 44.1 | $0.50 / $2.00 | 3 of 11 benchmarks | 55.4% | 1458 | 1340 | ||||||||
| 227 | Grok 3 Mini Betax-ai/grok-3-mini-beta | xAI | 43.9 | — | 1 of 11 benchmarks | 1386 | ||||||||||
| 228 | Mistral Large 3mistralai/mistral-large-3 | Mistral AI | 43.9 | — | 2 of 11 benchmarks | 1468 | 1230 | |||||||||
| 229 | Qwen3 30B A3Bqwen/qwen3-30b-a3b | Qwen | 43.8 | $0.12 / $0.50 | 1 of 11 benchmarks | 1386 | ||||||||||
| 230 | Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning | xAI | 43.7 | — | 2 of 11 benchmarks | 1461 | 1240 | |||||||||
| 231 | Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 43.4 | $0.25 / $1.50 | 2 of 11 benchmarks | 1457 | 1254 | ||||||||||
| 232 | QwQ 32Bqwen/qwq-32b | Qwen | 43.4 | — | 1 of 11 benchmarks | 1384 | ||||||||||
| 233 | GPT-5 Nano (high)openai/gpt-5-nano:high | OpenAI | 43.3 | $0.05 / $0.40 | 1 of 11 benchmarks | 1384 | ||||||||||
| 234 | Gemini 2.5 Flash Lite (thinking)google/gemini-2.5-flash-lite:thinking | 43.1 | $0.10 / $0.40 | 1 of 11 benchmarks | 1384 | |||||||||||
| 235 | OLMo 3.1 32B Instructallenai/olmo-3.1-32b-instruct | Allen Institute for AI | 43.0 | — | 1 of 11 benchmarks | 1382 | ||||||||||
| 236 | GPT-4.1 Nanoopenai/gpt-4.1-nano | OpenAI | 42.9 | $0.10 / $0.40 | 1 of 11 benchmarks | 1374 | ||||||||||
| 237 | Laguna XS.2poolside/laguna-xs.2 | Poolside | 42.8 | — | 1 of 11 benchmarks | 1303 | ||||||||||
| 238 | DeepSeek V4 Flash 0423 (xhigh)deepseek/deepseek-v4-flash:xhigh | DeepSeek | 42.7 | $0.14 / $0.28 | 2 of 11 benchmarks | 20.9% | 27.0 | |||||||||
| 239 | gpt-oss-20bopenai/gpt-oss-20b | OpenAI | 42.6 | $0.03 / $0.13 | 1 of 11 benchmarks | 1370 | ||||||||||
| 240 | Claude Haiku 4.5anthropic/claude-haiku-4.5 | Anthropic | 42.5 | $1.00 / $5.00 | 3 of 11 benchmarks | 64.7% | 1479 | 1326 | ||||||||
| 241 | Devstral Small 2507mistralai/devstral-small-2507 | Mistral AI | 42.5 | — | 1 of 11 benchmarks | 38.0% | ||||||||||
| 242 | Gemini 2.5 Progoogle/gemini-2.5-pro | 42.3 | $1.25 / $10.00 | 3 of 11 benchmarks | 57.6% | 1465 | 1226 | |||||||||
| 243 | Mercuryinception/mercury | Inception | 42.3 | — | 1 of 11 benchmarks | 1367 | ||||||||||
| 244 | GPT-5 Nano (medium)openai/gpt-5-nano:medium | OpenAI | 42.1 | $0.05 / $0.40 | 1 of 11 benchmarks | 34.8% | ||||||||||
| 245 | OLMo 3 32B Thinkallenai/olmo-3-32b-think | Allen Institute for AI | 42.0 | — | 1 of 11 benchmarks | 1364 | ||||||||||
| 246 | Llama 3.3 Nemotron Super 49B v1nvidia/llama-3.3-nemotron-super-49b-v1 | NVIDIA | 41.9 | — | 1 of 11 benchmarks | 1363 | ||||||||||
| 247 | Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | NVIDIA | 41.8 | $0.05 / $0.20 | 1 of 11 benchmarks | 1362 | ||||||||||
| 248 | GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | OpenAI | 41.7 | $2.50 / $10.00 | 1 of 11 benchmarks | 31.0% | ||||||||||
| 249 | Mistral Small 3.1 24B Instruct 2503mistralai/mistral-small-3.1-24b-instruct-2503 | Mistral AI | 41.5 | — | 1 of 11 benchmarks | 1362 | ||||||||||
| 250 | Gemma 3 27Bgoogle/gemma-3-27b-it | 41.2 | $0.08 / $0.45 | 1 of 11 benchmarks | 1358 | |||||||||||
| 251 | GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | OpenAI | 41.1 | $5.00 / $15.00 | 4 of 11 benchmarks | 38.8% | 31.3% | 12.0% | 1369 | |||||||
| 252 | Gemini 1.5 Pro 002google/gemini-1.5-pro-002 | 41.1 | — | 1 of 11 benchmarks | 1356 | |||||||||||
| 253 | KAT-Coder-Pro V1kwaipilot/kat-coder-pro-v1 | Kwaipilot | 40.9 | — | 1 of 11 benchmarks | 1255 | ||||||||||
| 254 | Hunyuan Large Visiontencent/hunyuan-large-vision | Tencent | 40.9 | — | 1 of 11 benchmarks | 1356 | ||||||||||
| 255 | Mimo v2 Flash (thinking)xiaomi/mimo-v2-flash:thinking | Xiaomi | 40.9 | — | 2 of 11 benchmarks | 1431 | 1293 | |||||||||
| 256 | Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct | Qwen | 40.8 | $0.36 / $0.40 | 1 of 11 benchmarks | 1356 | ||||||||||
| 257 | Mistral Large 2407mistralai/mistral-large-2407 | Mistral | 40.5 | $2.00 / $6.00 | 1 of 11 benchmarks | 1354 | ||||||||||
| 258 | Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | Qwen | 40.4 | $0.25 / $1.25 | 2 of 11 benchmarks | 1435 | 1250 | |||||||||
| 259 | Step 1o Turbo 202506stepfun/step-1o-turbo-202506 | StepFun | 40.2 | — | 1 of 11 benchmarks | 1352 | ||||||||||
| 260 | GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | OpenAI | 40.1 | $0.15 / $0.60 | 1 of 11 benchmarks | 1349 | ||||||||||
| 261 | GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini | OpenAI | 40.0 | $0.25 / $2.00 | 1 of 11 benchmarks | 1244 | ||||||||||
| 262 | GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | 40.0 | $0.40 / $1.60 | 2 of 11 benchmarks | 23.9% | 1433 | |||||||||
| 263 | GPT-4 Turboopenai/gpt-4-turbo | OpenAI | 40.0 | $10.00 / $30.00 | 1 of 11 benchmarks | 1347 | ||||||||||
| 264 | Qwen3.5-Flashqwen/qwen3.5-flash-02-23 | Qwen | 39.8 | $0.065 / $0.26 | 2 of 11 benchmarks | 1437 | 1238 | |||||||||
| 265 | Gemini 1.5 Pro 001google/gemini-1.5-pro-001 | 39.8 | — | 1 of 11 benchmarks | 1347 | |||||||||||
| 266 | DeepSeek V3.2deepseek/deepseek-v3.2 | DeepSeek | 39.8 | $0.269 / $0.40 | 3 of 11 benchmarks | 59.0% | 1470 | 1324 | ||||||||
| 267 | Mistral Large 2411mistralai/mistral-large-2411 | Mistral AI | 39.7 | — | 1 of 11 benchmarks | 1346 | ||||||||||
| 268 | Gemini 2.5 Flashgoogle/gemini-2.5-flash | 39.6 | $0.30 / $2.50 | 2 of 11 benchmarks | 28.7% | 1424 | ||||||||||
| 269 | GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | OpenAI | 39.6 | $2.50 / $10.00 | 4 of 11 benchmarks | 27.0% | 39.7% | 30.4% | 1360 | |||||||
| 270 | Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | Meta | 39.6 | $0.10 / $0.32 | 1 of 11 benchmarks | 1346 | ||||||||||
| 271 | Qwen 2.5qwen/qwen-2.5 | Qwen | 39.6 | — | 2 of 11 benchmarks | 40.2% | 24.7% | |||||||||
| 272 | Amazon Nova Pro v1.0amazon/amazon-nova-pro-v1.0 | Amazon | 39.4 | — | 1 of 11 benchmarks | 1343 | ||||||||||
| 273 | Gemini 2.0 Flash Lite Preview 02 05google/gemini-2.0-flash-lite-preview-02-05 | 39.3 | — | 1 of 11 benchmarks | 1343 | |||||||||||
| 274 | OLMo 3.1 32B Thinkallenai/olmo-3.1-32b-think | Allen Institute for AI | 38.9 | — | 1 of 11 benchmarks | 1338 | ||||||||||
| 275 | Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | Meta | 38.7 | $0.40 / $0.40 | 1 of 11 benchmarks | 1333 | ||||||||||
| 276 | Claude 3 Sonnetanthropic/claude-3-sonnet | Anthropic | 38.6 | — | 1 of 11 benchmarks | 1318 | ||||||||||
| 277 | Gemma 3 12Bgoogle/gemma-3-12b-it | 38.5 | $0.05 / $0.15 | 1 of 11 benchmarks | 1316 | |||||||||||
| 278 | MiniMax M2minimax/minimax-m2 | MiniMax | 38.4 | $0.255 / $1.02 | 3 of 11 benchmarks | 61.0% | 1385 | 1297 | ||||||||
| 279 | GPT 4 1106 Previewopenai/gpt-4-1106-preview | OpenAI | 38.4 | — | 4 of 11 benchmarks | 22.4% | 28.3% | 12.5% | 1340 | |||||||
| 280 | Gemini 1.5 Flash 002google/gemini-1.5-flash-002 | 38.3 | — | 1 of 11 benchmarks | 1316 | |||||||||||
| 281 | Mistral Small 24B Instruct 2501mistralai/mistral-small-24b-instruct-2501 | Mistral AI | 38.2 | — | 1 of 11 benchmarks | 1312 | ||||||||||
| 282 | Gemini 1.5 Flash 001google/gemini-1.5-flash-001 | 38.0 | — | 1 of 11 benchmarks | 1309 | |||||||||||
| 283 | Grok 4 Fast (reasoning)x-ai/grok-4-fast:reasoning | xAI | 37.9 | — | 2 of 11 benchmarks | 1436 | 1161 | |||||||||
| 284 | Gemma 3n E4Bgoogle/gemma-3n-e4b-it | 37.9 | — | 1 of 11 benchmarks | 1308 | |||||||||||
| 285 | Amazon Nova Lite v1.0amazon/amazon-nova-lite-v1.0 | Amazon | 37.8 | — | 1 of 11 benchmarks | 1306 | ||||||||||
| 286 | Trinity Large Thinkingarcee-ai/trinity-large-thinking | Arcee AI | 37.6 | $0.22 / $0.85 | 2 of 11 benchmarks | 1414 | 1239 | |||||||||
| 287 | Amazon Nova Micro v1.0amazon/amazon-nova-micro-v1.0 | Amazon | 37.5 | — | 1 of 11 benchmarks | 1289 | ||||||||||
| 288 | Command R (08-2024)cohere/command-r-08-2024 | Cohere | 37.4 | $0.15 / $0.60 | 1 of 11 benchmarks | 1281 | ||||||||||
| 289 | Command R+ (08-2024)cohere/command-r-plus-08-2024 | Cohere | 37.2 | $2.50 / $10.00 | 1 of 11 benchmarks | 1280 | ||||||||||
| 290 | OLMo 2 0325 32B Instructallenai/olmo-2-0325-32b-instruct | Allen Institute for AI | 37.1 | — | 1 of 11 benchmarks | 1280 | ||||||||||
| 291 | Grok Code Fast 1x-ai/grok-code-fast-1 | xAI | 37.0 | — | 1 of 11 benchmarks | 1164 | ||||||||||
| 292 | Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct | Mistral | 37.0 | $2.00 / $6.00 | 1 of 11 benchmarks | 1277 | ||||||||||
| 293 | Gemma 3 4Bgoogle/gemma-3-4b-it | 36.8 | — | 1 of 11 benchmarks | 1274 | |||||||||||
| 294 | gpt-oss-120bopenai/gpt-oss-120b | OpenAI | 36.7 | $0.03 / $0.17 | 2 of 11 benchmarks | 26.0% | 1390 | |||||||||
| 295 | Gemini 1.5 Flash 8B 001google/gemini-1.5-flash-8b-001 | 36.7 | — | 1 of 11 benchmarks | 1272 | |||||||||||
| 296 | Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | Meta | 36.5 | $0.05 / $0.08 | 1 of 11 benchmarks | 1260 | ||||||||||
| 297 | Gemini 3.1 Pro Preview (high)google/gemini-3.1-pro-preview:high | 36.4 | $2.00 / $12.00 | 1 of 11 benchmarks | 8.9% | |||||||||||
| 298 | Devstral Medium 2507mistralai/devstral-medium-2507 | Mistral AI | 36.4 | — | 1 of 11 benchmarks | 1080 | ||||||||||
| 299 | GPT-4oopenai/gpt-4o | OpenAI | 36.4 | $2.50 / $10.00 | 1 of 11 benchmarks | 12.2% | ||||||||||
| 300 | QwQ 32B Previewqwen/qwq-32b-preview | Qwen | 36.4 | — | 1 of 11 benchmarks | 1173 | ||||||||||
| 301 | Devstral 2mistralai/devstral-2 | Mistral AI | 36.1 | — | 2 of 11 benchmarks | 53.8% | 1194 | |||||||||
| 302 | Claude Opus 4.8 (medium)anthropic/claude-opus-4.8:medium | Anthropic | 35.9 | $5.00 / $25.00 | 2 of 11 benchmarks | 20.9% | 19.0 | |||||||||
| 303 | GPT-5 Miniopenai/gpt-5-mini | OpenAI | 35.8 | $0.25 / $2.00 | 2 of 11 benchmarks | 56.2% | 39.7% | |||||||||
| 304 | Mercury 2inception/mercury-2 | Inception | 34.6 | $0.25 / $0.75 | 2 of 11 benchmarks | 1394 | 1166 | |||||||||
| 305 | Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct | Meta | 34.3 | — | 2 of 11 benchmarks | 21.0% | 1373 | |||||||||
| 306 | Claude 3 Opusanthropic/claude-3-opus | Anthropic | 34.3 | — | 4 of 11 benchmarks | 15.8% | 26.3% | 10.5% | 1355 | |||||||
| 307 | Claude 3 Haikuanthropic/claude-3-haiku | Anthropic | 33.6 | $0.25 / $1.25 | 2 of 11 benchmarks | 40.6% | 1301 | |||||||||
| 308 | Gemini 2.0 Flash 001google/gemini-2.0-flash-001 | 33.3 | — | 2 of 11 benchmarks | 13.5% | 1365 | ||||||||||
| 309 | Claude 2anthropic/claude-2 | Anthropic | 32.8 | — | 3 of 11 benchmarks | 4.4% | 3.0% | 2.0% | ||||||||
| 310 | Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct | Meta | 32.6 | — | 2 of 11 benchmarks | 9.1% | 1362 | |||||||||
| 311 | Granite 4.1 8Bibm-granite/granite-4.1-8b | IBM | 31.2 | $0.05 / $0.10 | 2 of 11 benchmarks | 1353 | 1192 | |||||||||
| 312 | GLM 5.2 (high)z-ai/glm-5.2:high | Z.ai | 31.1 | $0.63 / $1.98 | 2 of 11 benchmarks | 17.4% | 18.0 | |||||||||
| 313 | Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct | Qwen | 30.5 | — | 2 of 11 benchmarks | 9.0% | 1342 | |||||||||
| 314 | SWE Llamaprinceton-nlp/swe-llama | Princeton NLP | 28.1 | — | 3 of 11 benchmarks | 1.4% | 1.3% | 0.7% | ||||||||
| 315 | Claude Opus 4.7 (medium)anthropic/claude-opus-4.7:medium | Anthropic | 27.3 | $5.00 / $25.00 | 2 of 11 benchmarks | 7.0% | 7.0 | |||||||||
| 316 | SWE Llama 13Bprinceton-nlp/swe-llama-13b | Princeton NLP | 26.5 | — | 3 of 11 benchmarks | 1.2% | 1.0% | 0.7% | ||||||||
| 317 | GPT 3.5openai/gpt-3.5 | OpenAI | 21.8 | — | 3 of 11 benchmarks | 0.4% | 0.3% | 0.2% |
How this ranks
Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.
A model scored on fewer than 3 of the 11 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.
Data sources
Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.
- Epoch AI Benchmarking HubCC BY 4.0
Benchmark runs by Epoch AI, from the AI Benchmarking Hub.
- LMArenaCC BY 4.0
Arena ratings by LMArena, from the public leaderboard dataset.
- SWE-bench
Resolve rates published by the SWE-bench maintainers.
- Warden
Security review results published by Warden.