Best LLM for math
Research-grade problem sets plus competition math. Contamination-resistant private sets carry the same weight as public ones.
Rumeqo runs these models inside your team rooms. See what each one costs.
| rank | model | vendor | composite | pricein / out | benchmarks | % | % | % | % | % | elo |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 (max)anthropic/claude-fable-5:max | Anthropic | 81.8 | $10.00 / $50.00 | 4 of 6 benchmarks | 100.0% | 87.0% | 100.0% | 99.7% | ||
| 2 | Claude Opus 5 (max)anthropic/claude-opus-5:max | Anthropic | 80.3 | $5.00 / $25.00 | 4 of 6 benchmarks | 85.6% | 73.2% | 98.9% | 1556 | ||
| 3 | GPT-5.6 Sol (max)openai/gpt-5.6-sol:max | OpenAI | 79.2 | $5.00 / $30.00 | 3 of 6 benchmarks | 89.1% | 82.9% | 100.0% | |||
| 4 | GPT-5.4 (high)openai/gpt-5.4:high | OpenAI | 78.7 | $2.50 / $15.00 | 4 of 6 benchmarks | 80.0% | 50.0% | 97.8% | 1494 | ||
| 5 | GPT-5.6 Terra (max)openai/gpt-5.6-terra:max | OpenAI | 76.6 | $1.00 / $6.00 | 3 of 6 benchmarks | 86.0% | 70.7% | 99.7% | |||
| 6 | Claude Opus 4.8 (max)anthropic/claude-opus-4.8:max | Anthropic | 74.3 | $5.00 / $25.00 | 4 of 6 benchmarks | 47.2% | 80.0% | 56.1% | 98.3% | ||
| 7 | GPT-5.6 Luna (max)openai/gpt-5.6-luna:max | OpenAI | 74.1 | $0.10 / $0.60 | 3 of 6 benchmarks | 82.1% | 61.0% | 98.3% | |||
| 8 | GPT 5.5 Pro Pre Release (xhigh)openai/gpt-5.5-pro-pre-release:xhigh | OpenAI | 73.9 | — | 3 of 6 benchmarks | 51.0% | 39.6% | 100.0% | |||
| 9 | GPT 5.5 Pre Release (xhigh)openai/gpt-5.5-pre-release:xhigh | OpenAI | 73.6 | — | 3 of 6 benchmarks | 51.7% | 35.4% | 100.0% | |||
| 10 | Kimi K3 (max)moonshotai/kimi-k3:max | MoonshotAI | 73.5 | $3.00 / $15.00 | 4 of 6 benchmarks | 72.2% | 39.0% | 97.2% | 1495 | ||
| 11 | GPT-5.5 Pro (xhigh)openai/gpt-5.5-pro:xhigh | OpenAI | 73.3 | $30.00 / $180.00 | 2 of 6 benchmarks | 87.7% | 78.0% | ||||
| 12 | Gemini 3.5 Flash (high)google/gemini-3.5-flash:high | 72.8 | $1.50 / $9.00 | 5 of 6 benchmarks | 80.0% | 62.8% | 26.8% | 95.6% | 1507 | ||
| 13 | Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 72.5 | $2.00 / $12.00 | 5 of 6 benchmarks | 88.9% | 59.6% | 26.8% | 95.6% | 1491 | ||
| 14 | GPT-5.4 (xhigh)openai/gpt-5.4:xhigh | OpenAI | 72.0 | $2.50 / $15.00 | 4 of 6 benchmarks | 47.6% | 78.6% | 49.0% | 95.3% | ||
| 15 | GPT-5.4 Pro (xhigh)openai/gpt-5.4-pro:xhigh | OpenAI | 72.0 | $30.00 / $180.00 | 3 of 6 benchmarks | 50.0% | 82.5% | 58.5% | |||
| 16 | Qwen3.8 Max (xhigh)qwen/qwen3.8-max:xhigh | Qwen | 71.8 | $2.00 / $6.00 | 3 of 6 benchmarks | 74.7% | 46.3% | 99.4% | |||
| 17 | GPT-5.2 (high)openai/gpt-5.2:high | OpenAI | 70.8 | $1.75 / $14.00 | 4 of 6 benchmarks | 60.0% | 18.8% | 96.1% | 1458 | ||
| 18 | GPT-5.5 (xhigh)openai/gpt-5.5:xhigh | OpenAI | 70.4 | $5.00 / $30.00 | 2 of 6 benchmarks | 85.3% | 72.5% | ||||
| 19 | Qwen3.7 Maxqwen/qwen3.7-max | Qwen | 69.9 | $1.475 / $4.425 | 4 of 6 benchmarks | 64.6% | 34.1% | 95.6% | 1489 | ||
| 20 | Claude Opus 4.6 (64K)anthropic/claude-opus-4.6:64k | Anthropic | 69.7 | $5.00 / $25.00 | 3 of 6 benchmarks | 90.0% | 20.8% | 94.4% | |||
| 21 | GPT-5.2 (xhigh)openai/gpt-5.2:xhigh | OpenAI | 68.7 | $1.75 / $14.00 | 4 of 6 benchmarks | 40.7% | 67.4% | 31.7% | 96.1% | ||
| 22 | Claude Opus 4.6 (max)anthropic/claude-opus-4.6:max | Anthropic | 68.6 | $5.00 / $25.00 | 4 of 6 benchmarks | 90.0% | 66.0% | 26.8% | 91.1% | ||
| 23 | GPT 5.5 Pro Pre Release (high)openai/gpt-5.5-pro-pre-release:high | OpenAI | 68.1 | — | 2 of 6 benchmarks | 52.4% | 39.6% | ||||
| 24 | Claude Opus 4.6 (32K)anthropic/claude-opus-4.6:32k | Anthropic | 68.0 | $5.00 / $25.00 | 3 of 6 benchmarks | 80.0% | 20.8% | 93.1% | |||
| 25 | DeepSeek V4 Pro (high)deepseek/deepseek-v4-pro:high | DeepSeek | 67.2 | $1.168 / $2.336 | 2 of 6 benchmarks | 95.6% | 1469 | ||||
| 26 | Kimi K2.6moonshotai/kimi-k2.6 | MoonshotAI | 66.9 | $0.95 / $4.00 | 5 of 6 benchmarks | 39.0% | 57.2% | 25.6% | 96.1% | 1478 | |
| 27 | GPT-5.2 (medium)openai/gpt-5.2:medium | OpenAI | 66.6 | $1.75 / $14.00 | 3 of 6 benchmarks | 60.0% | 16.7% | 93.9% | |||
| 28 | GLM 5.1z-ai/glm-5.1 | Z.ai | 66.3 | $1.40 / $4.40 | 4 of 6 benchmarks | 33.5% | 12.5% | 93.3% | 1481 | ||
| 29 | Gemini 3.6 Flash (high)google/gemini-3.6-flash:high | 66.3 | $1.50 / $7.50 | 4 of 6 benchmarks | 59.0% | 21.9% | 94.2% | 1513 | |||
| 30 | Qwen3.7 Plusqwen/qwen3.7-plus | Qwen | 66.0 | $0.32 / $1.28 | 2 of 6 benchmarks | 93.3% | 1471 | ||||
| 31 | Claude Fable 5anthropic/claude-fable-5 | Anthropic | 65.8 | $10.00 / $50.00 | 1 of 6 benchmarks | 1527 | |||||
| 32 | Claude Fable 5 (high)anthropic/claude-fable-5:high | Anthropic | 65.7 | $10.00 / $50.00 | 1 of 6 benchmarks | 100.0% | |||||
| 33 | Claude Opus 5 (high)anthropic/claude-opus-5:high | Anthropic | 65.7 | $5.00 / $25.00 | 1 of 6 benchmarks | 1525 | |||||
| 34 | Claude Opus 4.6 (high)anthropic/claude-opus-4.6:high | Anthropic | 65.5 | $5.00 / $25.00 | 1 of 6 benchmarks | 1516 | |||||
| 35 | Qwen3.8 Maxqwen/qwen3.8-max | Qwen | 65.4 | $2.00 / $6.00 | 1 of 6 benchmarks | 1513 | |||||
| 36 | Qwen3.6 Plusqwen/qwen3.6-plus | Qwen | 65.3 | $0.325 / $1.95 | 4 of 6 benchmarks | 50.0% | 8.3% | 93.3% | 1454 | ||
| 37 | Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | 65.3 | $5.00 / $25.00 | 3 of 6 benchmarks | 38.3% | 14.6% | 1506 | |||
| 38 | GPT-5 (high)openai/gpt-5:high | OpenAI | 65.0 | $1.25 / $10.00 | 6 of 6 benchmarks | 32.4% | 55.4% | 21.9% | 91.4% | 98.1% | 1435 |
| 39 | Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high | Anthropic | 64.8 | $5.00 / $25.00 | 1 of 6 benchmarks | 1503 | |||||
| 40 | GPT-5.5openai/gpt-5.5 | OpenAI | 64.7 | $5.00 / $30.00 | 1 of 6 benchmarks | 1499 | |||||
| 41 | Muse Sparkmeta/muse-spark | Meta | 64.6 | — | 4 of 6 benchmarks | 39.0% | 14.6% | 88.9% | 1461 | ||
| 42 | GPT-5.2 Pro (xhigh)openai/gpt-5.2-pro:xhigh | OpenAI | 64.6 | $21.00 / $168.00 | 2 of 6 benchmarks | 74.0% | 46.0% | ||||
| 43 | Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high | Anthropic | 64.2 | $5.00 / $25.00 | 1 of 6 benchmarks | 1493 | |||||
| 44 | Qwen3.6 Max Previewqwen/qwen3.6-max-preview | Qwen | 64.2 | $1.027 / $6.162 | 4 of 6 benchmarks | 50.0% | 4.2% | 91.1% | 1474 | ||
| 45 | GPT-5.5 (high)openai/gpt-5.5:high | OpenAI | 64.1 | $5.00 / $30.00 | 1 of 6 benchmarks | 1491 | |||||
| 46 | Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 64.0 | $0.50 / $3.00 | 5 of 6 benchmarks | 60.0% | 51.2% | 17.1% | 92.8% | 1476 | ||
| 47 | Claude Opus 4.7anthropic/claude-opus-4.7 | Anthropic | 63.9 | $5.00 / $25.00 | 1 of 6 benchmarks | 1491 | |||||
| 48 | Claude Fable 5 (low)anthropic/claude-fable-5:low | Anthropic | 63.8 | $10.00 / $50.00 | 1 of 6 benchmarks | 97.8% | |||||
| 49 | Claude Opus 4.8 (low)anthropic/claude-opus-4.8:low | Anthropic | 63.8 | $5.00 / $25.00 | 1 of 6 benchmarks | 97.8% | |||||
| 50 | Claude Opus 5anthropic/claude-opus-5 | Anthropic | 63.8 | $5.00 / $25.00 | 1 of 6 benchmarks | 97.8% | |||||
| 51 | GPT-5 (medium)openai/gpt-5:medium | OpenAI | 63.8 | $1.25 / $10.00 | 4 of 6 benchmarks | 27.2% | 6.3% | 87.2% | 97.9% | ||
| 52 | GLM 5.2 (max)z-ai/glm-5.2:max | Z.ai | 63.6 | $0.63 / $1.98 | 4 of 6 benchmarks | 59.2% | 29.3% | 86.4% | 1474 | ||
| 53 | Muse Spark 1.1meta/muse-spark-1.1 | Meta | 63.5 | $1.25 / $4.25 | 1 of 6 benchmarks | 1488 | |||||
| 54 | GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh | OpenAI | 63.4 | $0.10 / $0.60 | 1 of 6 benchmarks | 1484 | |||||
| 55 | Claude Opus 4.7 (max)anthropic/claude-opus-4.7:max | Anthropic | 63.2 | $5.00 / $25.00 | 3 of 6 benchmarks | 70.2% | 31.7% | 86.7% | |||
| 56 | GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh | OpenAI | 63.2 | $1.00 / $6.00 | 1 of 6 benchmarks | 1483 | |||||
| 57 | Hy3tencent/hy3 | Tencent | 63.1 | $0.132 / $0.528 | 1 of 6 benchmarks | 1481 | |||||
| 58 | Inklingthinkingmachines/inkling | Thinking Machines | 62.8 | $0.95 / $4.05 | 1 of 6 benchmarks | 1480 | |||||
| 59 | Grok 4.5x-ai/grok-4.5 | SpaceXAI | 62.7 | $2.00 / $6.00 | 1 of 6 benchmarks | 1479 | |||||
| 60 | GPT-5.1 (high)openai/gpt-5.1:high | OpenAI | 62.5 | $1.25 / $10.00 | 4 of 6 benchmarks | 31.0% | 12.5% | 88.6% | 1456 | ||
| 61 | Gemini 3 Progoogle/gemini-3-pro | 62.5 | — | 1 of 6 benchmarks | 1479 | ||||||
| 62 | Gemini 3 Pro Previewgoogle/gemini-3-pro-preview | 62.4 | — | 3 of 6 benchmarks | 37.6% | 18.8% | 91.4% | ||||
| 63 | GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh | OpenAI | 62.4 | $5.00 / $30.00 | 1 of 6 benchmarks | 1478 | |||||
| 64 | Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium | 62.1 | $1.50 / $9.00 | 1 of 6 benchmarks | 1478 | ||||||
| 65 | MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | Xiaomi | 62.0 | $0.435 / $0.87 | 1 of 6 benchmarks | 1477 | |||||
| 66 | Ernie 5.1baidu/ernie-5.1 | Baidu | 61.8 | — | 1 of 6 benchmarks | 1476 | |||||
| 67 | Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high | 61.8 | $0.50 / $3.00 | 1 of 6 benchmarks | 95.6% | ||||||
| 68 | Gemini 3.1 Pro Preview (high)google/gemini-3.1-pro-preview:high | 61.8 | $2.00 / $12.00 | 1 of 6 benchmarks | 95.6% | ||||||
| 69 | GPT-5.4 (medium)openai/gpt-5.4:medium | OpenAI | 61.8 | $2.50 / $15.00 | 1 of 6 benchmarks | 95.6% | |||||
| 70 | GPT-5.6 Sol (low)openai/gpt-5.6-sol:low | OpenAI | 61.8 | $5.00 / $30.00 | 1 of 6 benchmarks | 95.6% | |||||
| 71 | Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | Qwen | 61.7 | $0.50 / $3.60 | 2 of 6 benchmarks | 88.9% | 1449 | ||||
| 72 | Grok 4.5 (high)x-ai/grok-4.5:high | SpaceXAI | 61.3 | $2.00 / $6.00 | 3 of 6 benchmarks | 57.2% | 24.4% | 97.8% | |||
| 73 | Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | 61.2 | $5.00 / $25.00 | 1 of 6 benchmarks | 1472 | |||||
| 74 | Gemma 4 31Bgoogle/gemma-4-31b-it | 61.1 | $0.10 / $0.34 | 1 of 6 benchmarks | 1472 | ||||||
| 75 | Claude Sonnet 5 (xhigh)anthropic/claude-sonnet-5:xhigh | Anthropic | 60.9 | $2.00 / $10.00 | 1 of 6 benchmarks | 94.7% | |||||
| 76 | Qwen3.5 Max Previewqwen/qwen3.5-max-preview | Qwen | 60.8 | — | 1 of 6 benchmarks | 1470 | |||||
| 77 | Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06 | 60.8 | $1.25 / $10.00 | 1 of 6 benchmarks | 95.9% | ||||||
| 78 | Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking | MoonshotAI | 60.7 | $0.57 / $2.85 | 1 of 6 benchmarks | 1470 | |||||
| 79 | Claude Opus 4.5 (high 32K)anthropic/claude-opus-4.5:high-32k | Anthropic | 60.5 | $5.00 / $25.00 | 1 of 6 benchmarks | 1469 | |||||
| 80 | Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high | Anthropic | 60.2 | $2.00 / $10.00 | 1 of 6 benchmarks | 1468 | |||||
| 81 | Qwen3 Maxqwen/qwen3-max | Qwen | 60.2 | $0.78 / $3.90 | 3 of 6 benchmarks | 73.3% | 97.1% | 1426 | |||
| 82 | DeepSeek V4 Flash 0731 (max)deepseek/deepseek-v4-flash-0731:max | DeepSeek | 60.1 | $0.08 / $0.18 | 3 of 6 benchmarks | 57.5% | 24.4% | 94.4% | |||
| 83 | Gemma 4 26B A4B google/gemma-4-26b-a4b-it | 60.1 | $0.12 / $0.40 | 1 of 6 benchmarks | 1468 | ||||||
| 84 | Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning | xAI | 60.0 | — | 1 of 6 benchmarks | 1467 | |||||
| 85 | Claude Opus 5 (low)anthropic/claude-opus-5:low | Anthropic | 59.7 | $5.00 / $25.00 | 1 of 6 benchmarks | 93.3% | |||||
| 86 | Kimi K3 (high)moonshotai/kimi-k3:high | MoonshotAI | 59.7 | $3.00 / $15.00 | 1 of 6 benchmarks | 93.3% | |||||
| 87 | Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | 59.7 | $3.00 / $15.00 | 1 of 6 benchmarks | 1463 | |||||
| 88 | GPT 5.5 Instantopenai/gpt-5.5-instant | OpenAI | 59.5 | — | 1 of 6 benchmarks | 1462 | |||||
| 89 | GPT-5.4openai/gpt-5.4 | OpenAI | 59.2 | $2.50 / $15.00 | 1 of 6 benchmarks | 1460 | |||||
| 90 | GPT-5 Pro (high)openai/gpt-5-pro:high | OpenAI | 58.9 | $15.00 / $120.00 | 3 of 6 benchmarks | 60.0% | 55.8% | 19.5% | |||
| 91 | Claude Sonnet 4.5 (high 32K)anthropic/claude-sonnet-4.5:high-32k | Anthropic | 58.8 | $3.00 / $15.00 | 1 of 6 benchmarks | 1455 | |||||
| 92 | Claude Sonnet 5 (max)anthropic/claude-sonnet-5:max | Anthropic | 58.7 | $2.00 / $10.00 | 3 of 6 benchmarks | 65.6% | 29.3% | 80.0% | |||
| 93 | Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal | 58.5 | $0.50 / $3.00 | 1 of 6 benchmarks | 1453 | ||||||
| 94 | Muse Glimmer 30Bmeta/muse-glimmer-30b | Meta | 58.3 | $0.35 / $1.50 | 1 of 6 benchmarks | 1453 | |||||
| 95 | Grok 4.20 Multi Agent Beta 0309x-ai/grok-4.20-multi-agent-beta-0309 | xAI | 58.3 | — | 1 of 6 benchmarks | 1453 | |||||
| 96 | Mimo v2 Proxiaomi/mimo-v2-pro | Xiaomi | 58.1 | — | 1 of 6 benchmarks | 1452 | |||||
| 97 | Qwen3.6 27Bqwen/qwen3.6-27b | Qwen | 58.0 | $0.60 / $3.60 | 1 of 6 benchmarks | 91.1% | |||||
| 98 | GPT-5.2 Chatopenai/gpt-5.2-chat | OpenAI | 58.0 | $1.75 / $14.00 | 1 of 6 benchmarks | 1452 | |||||
| 99 | GPT-5.4 Mini (high)openai/gpt-5.4-mini:high | OpenAI | 57.8 | $0.75 / $4.50 | 4 of 6 benchmarks | 50.0% | 2.1% | 87.2% | 1439 | ||
| 100 | Kimi K2p5moonshotai/kimi-k2p5 | Moonshot AI | 57.8 | — | 3 of 6 benchmarks | 27.9% | 4.2% | 92.2% | |||
| 101 | Dola Seed 2.0 Probytedance/dola-seed-2.0-pro | ByteDance | 57.8 | — | 1 of 6 benchmarks | 1451 | |||||
| 102 | o3openai/o3 | OpenAI | 57.5 | $2.00 / $8.00 | 1 of 6 benchmarks | 1447 | |||||
| 103 | GPT-5 Mini (high)openai/gpt-5-mini:high | OpenAI | 57.5 | $0.25 / $2.00 | 6 of 6 benchmarks | 27.2% | 46.7% | 12.2% | 86.7% | 97.8% | 1405 |
| 104 | DeepSeek V4 Prodeepseek/deepseek-v4-pro | DeepSeek | 57.4 | $1.168 / $2.336 | 1 of 6 benchmarks | 1445 | |||||
| 105 | Claude Opus 4.1 (thinking 16K)anthropic/claude-opus-4.1:thinking-16k | Anthropic | 57.2 | $15.00 / $75.00 | 1 of 6 benchmarks | 1444 | |||||
| 106 | Gemini 3.5 Flash (low)google/gemini-3.5-flash:low | 57.2 | $1.50 / $9.00 | 1 of 6 benchmarks | 88.9% | ||||||
| 107 | GPT-5.6 Terra (low)openai/gpt-5.6-terra:low | OpenAI | 57.2 | $1.00 / $6.00 | 1 of 6 benchmarks | 88.9% | |||||
| 108 | gpt-oss-120b (high)openai/gpt-oss-120b:high | OpenAI | 57.2 | $0.03 / $0.17 | 1 of 6 benchmarks | 88.9% | |||||
| 109 | GLM 5V Turboz-ai/glm-5v-turbo | Z.ai | 57.1 | $1.20 / $4.00 | 1 of 6 benchmarks | 1443 | |||||
| 110 | GPT-5 Mini (medium)openai/gpt-5-mini:medium | OpenAI | 57.0 | $0.25 / $2.00 | 4 of 6 benchmarks | 20.3% | 4.2% | 78.3% | 96.8% | ||
| 111 | Grok 4.1 (thinking)x-ai/grok-4.1:thinking | xAI | 56.8 | — | 1 of 6 benchmarks | 1442 | |||||
| 112 | Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | NVIDIA | 56.7 | $0.60 / $3.60 | 1 of 6 benchmarks | 1442 | |||||
| 113 | Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo | Moonshot AI | 56.6 | — | 3 of 6 benchmarks | 20.0% | 83.1% | 1436 | |||
| 114 | MiMo-V2.5xiaomi/mimo-v2.5 | Xiaomi | 56.5 | $0.14 / $0.28 | 1 of 6 benchmarks | 1441 | |||||
| 115 | Claude Opus 4.7 (xhigh)anthropic/claude-opus-4.7:xhigh | Anthropic | 56.3 | $5.00 / $25.00 | 3 of 6 benchmarks | 43.8% | 0.0% | 97.8% | |||
| 116 | DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high | DeepSeek | 56.2 | $0.14 / $0.28 | 1 of 6 benchmarks | 1441 | |||||
| 117 | Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant | Moonshot AI | 55.9 | — | 1 of 6 benchmarks | 1440 | |||||
| 118 | o1 (medium)openai/o1:medium | OpenAI | 55.8 | $15.00 / $60.00 | 2 of 6 benchmarks | 73.3% | 94.4% | ||||
| 119 | GPT-5.2 (low)openai/gpt-5.2:low | OpenAI | 55.7 | $1.75 / $14.00 | 3 of 6 benchmarks | 40.0% | 6.3% | 78.9% | |||
| 120 | Ernie 5.0 0110baidu/ernie-5.0-0110 | Baidu | 55.7 | — | 1 of 6 benchmarks | 1437 | |||||
| 121 | Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | Qwen | 55.6 | $0.15 / $1.00 | 1 of 6 benchmarks | 86.7% | |||||
| 122 | Qwen3.7 Flashqwen/qwen3.7-flash | Qwen | 55.6 | $0.03 / $0.13 | 1 of 6 benchmarks | 86.7% | |||||
| 123 | Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 55.5 | $0.25 / $1.50 | 1 of 6 benchmarks | 1437 | ||||||
| 124 | o3 (high)openai/o3:high | OpenAI | 55.1 | $2.00 / $8.00 | 4 of 6 benchmarks | 18.7% | 2.1% | 83.9% | 97.8% | ||
| 125 | Qwen3.5 Plusqwen/qwen3.5-plus | Qwen | 55.1 | $0.30 / $1.80 | 3 of 6 benchmarks | 50.0% | 2.1% | 86.7% | |||
| 126 | Mimo v2 Omnixiaomi/mimo-v2-omni | Xiaomi | 55.1 | — | 1 of 6 benchmarks | 1435 | |||||
| 127 | GPT-5.2openai/gpt-5.2 | OpenAI | 54.8 | $1.75 / $14.00 | 1 of 6 benchmarks | 1432 | |||||
| 128 | Claude Sonnet 4.6 (32K)anthropic/claude-sonnet-4.6:32k | Anthropic | 54.7 | $3.00 / $15.00 | 1 of 6 benchmarks | 85.8% | |||||
| 129 | Kimi K2.7 Codemoonshotai/kimi-k2.7-code | MoonshotAI | 54.6 | $0.67 / $3.40 | 3 of 6 benchmarks | 54.0% | 12.2% | 95.6% | |||
| 130 | Mistral Medium 3.5mistralai/mistral-medium-3-5 | Mistral | 54.6 | $1.50 / $7.50 | 1 of 6 benchmarks | 1429 | |||||
| 131 | DeepSeek V3.2 Exp (thinking)deepseek/deepseek-v3.2-exp:thinking | DeepSeek | 54.5 | $0.27 / $0.41 | 1 of 6 benchmarks | 1429 | |||||
| 132 | Grok 4.1x-ai/grok-4.1 | xAI | 54.2 | — | 1 of 6 benchmarks | 1428 | |||||
| 133 | Qwen3.5-27Bqwen/qwen3.5-27b | Qwen | 54.1 | $0.195 / $1.56 | 1 of 6 benchmarks | 1428 | |||||
| 134 | GPT-5.1 (medium)openai/gpt-5.1:medium | OpenAI | 54.0 | $1.25 / $10.00 | 3 of 6 benchmarks | 26.9% | 4.2% | 85.6% | |||
| 135 | o4 Mini Highopenai/o4-mini-high | OpenAI | 53.9 | $1.10 / $4.40 | 5 of 6 benchmarks | 24.8% | 36.1% | 4.9% | 81.7% | 97.8% | |
| 136 | Claude Opus 4.8 (none)anthropic/claude-opus-4.8:none | Anthropic | 53.8 | $5.00 / $25.00 | 1 of 6 benchmarks | 84.4% | |||||
| 137 | GPT-5.4 (low)openai/gpt-5.4:low | OpenAI | 53.8 | $2.50 / $15.00 | 1 of 6 benchmarks | 84.4% | |||||
| 138 | GPT-5.5 (low)openai/gpt-5.5:low | OpenAI | 53.8 | $5.00 / $30.00 | 1 of 6 benchmarks | 84.4% | |||||
| 139 | MiniMax M3minimax/minimax-m3 | MiniMax | 53.8 | $0.30 / $1.20 | 2 of 6 benchmarks | 71.1% | 1440 | ||||
| 140 | GPT-5.3 Chatopenai/gpt-5.3-chat | OpenAI | 53.5 | — | 1 of 6 benchmarks | 1427 | |||||
| 141 | DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | DeepSeek | 53.2 | $0.14 / $0.28 | 1 of 6 benchmarks | 1425 | |||||
| 142 | Hunyuan Hy3 Previewtencent/hunyuan-hy3-preview | Tencent | 53.1 | — | 1 of 6 benchmarks | 1425 | |||||
| 143 | DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking | DeepSeek | 52.9 | $0.269 / $0.40 | 1 of 6 benchmarks | 1424 | |||||
| 144 | o3 (medium)openai/o3:medium | OpenAI | 52.9 | $2.00 / $8.00 | 2 of 6 benchmarks | 16.9% | 84.4% | ||||
| 145 | Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 52.8 | $0.30 / $2.50 | 1 of 6 benchmarks | 1423 | ||||||
| 146 | Claude Sonnet 4.6 (medium)anthropic/claude-sonnet-4.6:medium | Anthropic | 52.6 | $3.00 / $15.00 | 1 of 6 benchmarks | 82.2% | |||||
| 147 | Gemini 3.6 Flash (low)google/gemini-3.6-flash:low | 52.6 | $1.50 / $7.50 | 1 of 6 benchmarks | 82.2% | ||||||
| 148 | Qwen3.5 397B A17B (none)qwen/qwen3.5-397b-a17b:none | Qwen | 52.6 | $0.50 / $3.60 | 1 of 6 benchmarks | 82.2% | |||||
| 149 | GPT-5.4 Nano (high)openai/gpt-5.4-nano:high | OpenAI | 52.6 | $0.20 / $1.25 | 5 of 6 benchmarks | 25.9% | 44.9% | 12.2% | 87.8% | 1423 | |
| 150 | o1 (high)openai/o1:high | OpenAI | 52.6 | $15.00 / $60.00 | 3 of 6 benchmarks | 9.3% | 73.3% | 94.7% | |||
| 151 | Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | Qwen | 52.5 | $0.29 / $2.40 | 1 of 6 benchmarks | 1423 | |||||
| 152 | GPT-5.1openai/gpt-5.1 | OpenAI | 52.4 | $1.25 / $10.00 | 1 of 6 benchmarks | 1423 | |||||
| 153 | o3 Mini (medium)openai/o3-mini:medium | OpenAI | 52.3 | $1.10 / $4.40 | 3 of 6 benchmarks | 11.3% | 63.9% | 95.2% | |||
| 154 | MiniMax M2.7minimax/minimax-m2.7 | MiniMax | 52.2 | $0.30 / $1.20 | 1 of 6 benchmarks | 1423 | |||||
| 155 | Claude Opus 4 (thinking 16K)anthropic/claude-opus-4:thinking-16k | Anthropic | 52.1 | $15.00 / $75.00 | 1 of 6 benchmarks | 1421 | |||||
| 156 | Grok 4.20 0309 (reasoning)x-ai/grok-4.20-0309:reasoning | xAI | 51.9 | — | 3 of 6 benchmarks | 44.9% | 17.1% | 92.2% | |||
| 157 | Grok 4.3x-ai/grok-4.3 | SpaceXAI | 51.8 | $1.25 / $2.50 | 1 of 6 benchmarks | 1419 | |||||
| 158 | Claude Opus 4.5 (16K)anthropic/claude-opus-4.5:16k | Anthropic | 51.6 | $5.00 / $25.00 | 3 of 6 benchmarks | 40.0% | 2.1% | 81.7% | |||
| 159 | Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | Qwen | 51.6 | $0.09 / $0.55 | 1 of 6 benchmarks | 1418 | |||||
| 160 | Claude Opus 4.5anthropic/claude-opus-4.5 | Anthropic | 51.5 | $5.00 / $25.00 | 4 of 6 benchmarks | 20.7% | 4.2% | 48.1% | 1465 | ||
| 161 | Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning | xAI | 51.5 | — | 1 of 6 benchmarks | 1418 | |||||
| 162 | Gemini 3.1 Flash Lite (high)google/gemini-3.1-flash-lite:high | 51.4 | $0.25 / $1.50 | 1 of 6 benchmarks | 80.0% | ||||||
| 163 | Gemini 3.5 Flash (minimal)google/gemini-3.5-flash:minimal | 51.4 | $1.50 / $9.00 | 1 of 6 benchmarks | 80.0% | ||||||
| 164 | Gemini 3.6 Flash (minimal)google/gemini-3.6-flash:minimal | 51.4 | $1.50 / $7.50 | 1 of 6 benchmarks | 80.0% | ||||||
| 165 | Qwen3.7 Plus (none)qwen/qwen3.7-plus:none | Qwen | 51.4 | $0.32 / $1.28 | 1 of 6 benchmarks | 80.0% | |||||
| 166 | Kimi K2 0905moonshotai/kimi-k2-0905 | MoonshotAI | 51.4 | $0.60 / $2.50 | 1 of 6 benchmarks | 1417 | |||||
| 167 | DeepSeek V3.2deepseek/deepseek-v3.2 | DeepSeek | 51.3 | $0.269 / $0.40 | 3 of 6 benchmarks | 22.1% | 2.1% | 1429 | |||
| 168 | Grok 4.3 (high)x-ai/grok-4.3:high | SpaceXAI | 51.2 | $1.25 / $2.50 | 3 of 6 benchmarks | 42.8% | 14.6% | 93.3% | |||
| 169 | R1deepseek/deepseek-r1 | DeepSeek | 51.2 | $0.70 / $2.50 | 3 of 6 benchmarks | 53.3% | 93.0% | 1412 | |||
| 170 | Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct | Qwen | 51.2 | $0.10 / $1.10 | 1 of 6 benchmarks | 1417 | |||||
| 171 | GLM 5z-ai/glm-5 | Z.ai | 51.2 | $0.95 / $2.55 | 4 of 6 benchmarks | 16.4% | 2.1% | 80.0% | 1443 | ||
| 172 | DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | DeepSeek | 51.1 | $0.27 / $0.41 | 1 of 6 benchmarks | 1416 | |||||
| 173 | o4 Miniopenai/o4-mini | OpenAI | 50.9 | $1.10 / $4.40 | 1 of 6 benchmarks | 1415 | |||||
| 174 | LongCat Flash Chatmeituan/longcat-flash-chat | Meituan | 50.8 | — | 1 of 6 benchmarks | 1415 | |||||
| 175 | GPT-5.4 Mini (xhigh)openai/gpt-5.4-mini:xhigh | OpenAI | 50.7 | $0.75 / $4.50 | 3 of 6 benchmarks | 51.2% | 9.8% | 88.9% | |||
| 176 | DeepSeek V3.1 (thinking)deepseek/deepseek-chat-v3.1:thinking | DeepSeek | 50.6 | $0.25 / $0.95 | 1 of 6 benchmarks | 1414 | |||||
| 177 | Claude Sonnet 4.6 (16K)anthropic/claude-sonnet-4.6:16k | Anthropic | 50.5 | $3.00 / $15.00 | 2 of 6 benchmarks | 80.0% | 0.0% | ||||
| 178 | DeepSeek V3.1deepseek/deepseek-chat-v3.1 | DeepSeek | 50.5 | $0.25 / $0.95 | 1 of 6 benchmarks | 1413 | |||||
| 179 | Qwen3.7 Flash (none)qwen/qwen3.7-flash:none | Qwen | 50.3 | $0.03 / $0.13 | 1 of 6 benchmarks | 77.8% | |||||
| 180 | GPT 5 Chatopenai/gpt-5-chat | OpenAI | 50.2 | — | 1 of 6 benchmarks | 1413 | |||||
| 181 | Grok 4 Fast (reasoning)x-ai/grok-4-fast:reasoning | xAI | 49.9 | — | 1 of 6 benchmarks | 1410 | |||||
| 182 | DeepSeek V4 Pro (max)deepseek/deepseek-v4-pro:max | DeepSeek | 49.9 | $1.168 / $2.336 | 3 of 6 benchmarks | 45.3% | 2.4% | 96.7% | |||
| 183 | Ernie 5.0 Preview 1203baidu/ernie-5.0-preview-1203 | Baidu | 49.8 | — | 1 of 6 benchmarks | 1409 | |||||
| 184 | Claude Sonnet 4.6 (high)anthropic/claude-sonnet-4.6:high | Anthropic | 49.7 | $3.00 / $15.00 | 1 of 6 benchmarks | 75.6% | |||||
| 185 | Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct | Qwen | 49.5 | $0.26 / $1.04 | 1 of 6 benchmarks | 1408 | |||||
| 186 | GPT-5 Nano (medium)openai/gpt-5-nano:medium | OpenAI | 49.4 | $0.05 / $0.40 | 4 of 6 benchmarks | 10.0% | 2.1% | 74.2% | 95.2% | ||
| 187 | Claude Sonnet 4.5 (59K)anthropic/claude-sonnet-4.5:59k | Anthropic | 49.3 | $3.00 / $15.00 | 2 of 6 benchmarks | 13.5% | 77.8% | ||||
| 188 | Step 3.5 Flashstepfun/step-3.5-flash | StepFun | 49.2 | $0.10 / $0.30 | 1 of 6 benchmarks | 1406 | |||||
| 189 | Gemma 4 31B (minimal)google/gemma-4-31b-it:minimal | 49.0 | $0.10 / $0.34 | 1 of 6 benchmarks | 73.3% | ||||||
| 190 | Mistral Large 3mistralai/mistral-large-3 | Mistral AI | 48.6 | — | 1 of 6 benchmarks | 1404 | |||||
| 191 | Chatgpt 4oopenai/chatgpt-4o | OpenAI | 48.5 | — | 1 of 6 benchmarks | 1404 | |||||
| 192 | Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking | Qwen | 48.4 | $0.40 / $4.00 | 1 of 6 benchmarks | 1404 | |||||
| 193 | Claude Sonnet 4 (32K)anthropic/claude-sonnet-4:32k | Anthropic | 48.2 | $3.00 / $15.00 | 1 of 6 benchmarks | 71.1% | |||||
| 194 | Claude Sonnet 4.5 (16K)anthropic/claude-sonnet-4.5:16k | Anthropic | 48.2 | $3.00 / $15.00 | 1 of 6 benchmarks | 71.1% | |||||
| 195 | Claude Sonnet 4.6 (max)anthropic/claude-sonnet-4.6:max | Anthropic | 48.2 | $3.00 / $15.00 | 1 of 6 benchmarks | 71.1% | |||||
| 196 | Gemini 3.5 Flash Lite (high)google/gemini-3.5-flash-lite:high | 48.2 | $0.30 / $2.50 | 1 of 6 benchmarks | 71.1% | ||||||
| 197 | Gemini 2.5 Progoogle/gemini-2.5-pro | 48.0 | $1.25 / $10.00 | 6 of 6 benchmarks | 14.1% | 24.6% | 0.0% | 84.7% | 95.6% | 1441 | |
| 198 | Grok 4 0709x-ai/grok-4-0709 | xAI | 48.0 | — | 3 of 6 benchmarks | 19.7% | 2.1% | 1427 | |||
| 199 | Claude Sonnet 4 (thinking 32K)anthropic/claude-sonnet-4:thinking-32k | Anthropic | 47.9 | $3.00 / $15.00 | 1 of 6 benchmarks | 1403 | |||||
| 200 | Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-preview | 47.9 | $1.25 / $10.00 | 2 of 6 benchmarks | 30.0% | 2.1% | |||||
| 201 | Hunyuan T1tencent/hunyuan-t1 | Tencent | 47.8 | — | 1 of 6 benchmarks | 1401 | |||||
| 202 | Claude Sonnet 4.5 (32K)anthropic/claude-sonnet-4.5:32k | Anthropic | 47.8 | $3.00 / $15.00 | 5 of 6 benchmarks | 15.2% | 23.9% | 2.4% | 77.8% | 97.7% | |
| 203 | Claude Haiku 4.5 (32K)anthropic/claude-haiku-4.5:32k | Anthropic | 47.6 | $1.00 / $5.00 | 4 of 6 benchmarks | 5.9% | 2.1% | 66.7% | 96.4% | ||
| 204 | Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | Qwen | 47.6 | $0.25 / $1.25 | 1 of 6 benchmarks | 1400 | |||||
| 205 | Qwen3 32Bqwen/qwen3-32b | Qwen | 47.5 | $0.08 / $0.28 | 1 of 6 benchmarks | 1399 | |||||
| 206 | Claude Opus 4.5 (32K)anthropic/claude-opus-4.5:32k | Anthropic | 47.3 | $5.00 / $25.00 | 4 of 6 benchmarks | 20.7% | 34.4% | 4.9% | 86.1% | ||
| 207 | Mistral Medium 2508mistralai/mistral-medium-2508 | Mistral AI | 47.2 | — | 1 of 6 benchmarks | 1398 | |||||
| 208 | Kimi K3 (low)moonshotai/kimi-k3:low | MoonshotAI | 47.2 | $3.00 / $15.00 | 1 of 6 benchmarks | 68.9% | |||||
| 209 | GPT-5.4 Nano (low)openai/gpt-5.4-nano:low | OpenAI | 47.2 | $0.20 / $1.25 | 1 of 6 benchmarks | 68.9% | |||||
| 210 | GPT-5.6 Sol (none)openai/gpt-5.6-sol:none | OpenAI | 47.2 | $5.00 / $30.00 | 1 of 6 benchmarks | 68.9% | |||||
| 211 | Qwen3.6 35B A3B (none)qwen/qwen3.6-35b-a3b:none | Qwen | 47.2 | $0.15 / $1.00 | 1 of 6 benchmarks | 68.9% | |||||
| 212 | Ernie 5.0 Preview 1022baidu/ernie-5.0-preview-1022 | Baidu | 46.9 | — | 1 of 6 benchmarks | 1396 | |||||
| 213 | GPT-5.1 (low)openai/gpt-5.1:low | OpenAI | 46.9 | $1.25 / $10.00 | 2 of 6 benchmarks | 17.3% | 63.9% | ||||
| 214 | MiniMax M2.5minimax/minimax-m2.5 | MiniMax | 46.8 | $0.22 / $0.90 | 1 of 6 benchmarks | 1396 | |||||
| 215 | DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | DeepSeek | 46.5 | $0.27 / $0.95 | 1 of 6 benchmarks | 1394 | |||||
| 216 | R1 0528deepseek/deepseek-r1-0528 | DeepSeek | 46.4 | $0.50 / $2.15 | 4 of 6 benchmarks | 0.0% | 66.4% | 96.6% | 1395 | ||
| 217 | GPT-5.6 Luna (low)openai/gpt-5.6-luna:low | OpenAI | 46.4 | $0.10 / $0.60 | 1 of 6 benchmarks | 66.7% | |||||
| 218 | Qwen3.6 27B (none)qwen/qwen3.6-27b:none | Qwen | 46.4 | $0.60 / $3.60 | 1 of 6 benchmarks | 66.7% | |||||
| 219 | Qwen3 235B A22B (nothinking)qwen/qwen3-235b-a22b:nothinking | Qwen | 46.4 | $0.455 / $1.82 | 1 of 6 benchmarks | 1393 | |||||
| 220 | Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | Qwen | 46.1 | $0.15 / $1.20 | 1 of 6 benchmarks | 1391 | |||||
| 221 | Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | Qwen | 46.0 | $0.23 / $2.30 | 4 of 6 benchmarks | 20.0% | 0.0% | 86.7% | 1397 | ||
| 222 | MiniMax M2.1minimax/minimax-m2.1 | MiniMax | 45.9 | $0.30 / $1.20 | 1 of 6 benchmarks | 1390 | |||||
| 223 | Grok 3 Mini Beta (low)x-ai/grok-3-mini-beta:low | xAI | 45.8 | — | 3 of 6 benchmarks | 2.8% | 62.2% | 90.9% | |||
| 224 | GLM 4.5 Airz-ai/glm-4.5-air | Z.ai | 45.8 | $0.13 / $0.85 | 1 of 6 benchmarks | 1390 | |||||
| 225 | Kimi K2 0711moonshotai/kimi-k2 | MoonshotAI | 45.5 | $0.57 / $2.30 | 1 of 6 benchmarks | 1388 | |||||
| 226 | Qwen3.6 Flashqwen/qwen3.6-flash | Qwen | 45.4 | $0.1875 / $1.125 | 3 of 6 benchmarks | 20.0% | 0.0% | 84.4% | |||
| 227 | GPT-5.2 (none)openai/gpt-5.2:none | OpenAI | 45.2 | $1.75 / $14.00 | 1 of 6 benchmarks | 62.2% | |||||
| 228 | Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | NVIDIA | 45.2 | — | 1 of 6 benchmarks | 1386 | |||||
| 229 | Claude 3.7 Sonnet (thinking 32K)anthropic/claude-3-7-sonnet:thinking-32k | Anthropic | 45.1 | — | 1 of 6 benchmarks | 1385 | |||||
| 230 | o4 Mini (medium)openai/o4-mini:medium | OpenAI | 44.9 | $1.10 / $4.40 | 3 of 6 benchmarks | 19.0% | 2.1% | 73.3% | |||
| 231 | Trinity Large Thinkingarcee-ai/trinity-large-thinking | Arcee AI | 44.9 | $0.22 / $0.85 | 1 of 6 benchmarks | 1384 | |||||
| 232 | Claude Opus 4 (16K)anthropic/claude-opus-4:16k | Anthropic | 44.8 | $15.00 / $75.00 | 1 of 6 benchmarks | 60.0% | |||||
| 233 | Gemini 3.5 Flash Lite (low)google/gemini-3.5-flash-lite:low | 44.8 | $0.30 / $2.50 | 1 of 6 benchmarks | 60.0% | ||||||
| 234 | Claude 3.7 Sonnet (32K)anthropic/claude-3-7-sonnet:32k | Anthropic | 44.8 | — | 3 of 6 benchmarks | 3.5% | 53.3% | 90.0% | |||
| 235 | INTELLECT-3prime-intellect/intellect-3 | Prime Intellect | 44.8 | — | 1 of 6 benchmarks | 1382 | |||||
| 236 | o3 Miniopenai/o3-mini | OpenAI | 44.6 | $1.10 / $4.40 | 1 of 6 benchmarks | 1382 | |||||
| 237 | gpt-oss-120bopenai/gpt-oss-120b | OpenAI | 44.5 | $0.03 / $0.17 | 1 of 6 benchmarks | 1381 | |||||
| 238 | Claude Opus 4.1anthropic/claude-opus-4.1 | Anthropic | 44.4 | $15.00 / $75.00 | 3 of 6 benchmarks | 5.9% | 40.0% | 1434 | |||
| 239 | Llama 3.1 Nemotron Ultra 253B v1nvidia/llama-3.1-nemotron-ultra-253b-v1 | NVIDIA | 44.4 | — | 1 of 6 benchmarks | 1380 | |||||
| 240 | Gemini 2.5 Flashgoogle/gemini-2.5-flash | 44.3 | $0.30 / $2.50 | 4 of 6 benchmarks | 4.8% | 4.2% | 70.8% | 1406 | |||
| 241 | GPT-5.4 (none)openai/gpt-5.4:none | OpenAI | 44.3 | $2.50 / $15.00 | 1 of 6 benchmarks | 57.8% | |||||
| 242 | GPT-5.5 (none)openai/gpt-5.5:none | OpenAI | 44.3 | $5.00 / $30.00 | 1 of 6 benchmarks | 57.8% | |||||
| 243 | Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | Qwen | 44.2 | $0.0482 / $0.1931 | 1 of 6 benchmarks | 1380 | |||||
| 244 | GPT-5 Nano (high)openai/gpt-5-nano:high | OpenAI | 44.2 | $0.05 / $0.40 | 6 of 6 benchmarks | 20.0% | 20.0% | 2.4% | 81.1% | 94.9% | 1346 |
| 245 | Mimo v2 Flashxiaomi/mimo-v2-flash | Xiaomi | 44.1 | — | 1 of 6 benchmarks | 1377 | |||||
| 246 | o4 Mini (low)openai/o4-mini:low | OpenAI | 44.0 | $1.10 / $4.40 | 2 of 6 benchmarks | 10.7% | 57.8% | ||||
| 247 | o1openai/o1 | OpenAI | 44.0 | $15.00 / $60.00 | 3 of 6 benchmarks | 31.1% | 81.7% | 1409 | |||
| 248 | Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | NVIDIA | 43.9 | $0.085 / $0.40 | 1 of 6 benchmarks | 1376 | |||||
| 249 | GPT 4.5 Previewopenai/gpt-4.5-preview | OpenAI | 43.9 | — | 3 of 6 benchmarks | 37.8% | 78.6% | 1408 | |||
| 250 | Claude Opus 4.1 (27K)anthropic/claude-opus-4.1:27k | Anthropic | 43.9 | $15.00 / $75.00 | 3 of 6 benchmarks | 7.2% | 4.2% | 68.9% | |||
| 251 | Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct | Qwen | 43.8 | — | 1 of 6 benchmarks | 1376 | |||||
| 252 | GPT-5 Mini (minimal)openai/gpt-5-mini:minimal | OpenAI | 43.8 | $0.25 / $2.00 | 1 of 6 benchmarks | 55.6% | |||||
| 253 | Mimo v2 Flash (thinking)xiaomi/mimo-v2-flash:thinking | Xiaomi | 43.6 | — | 1 of 6 benchmarks | 1374 | |||||
| 254 | o3 (low)openai/o3:low | OpenAI | 43.5 | $2.00 / $8.00 | 2 of 6 benchmarks | 9.7% | 60.0% | ||||
| 255 | gpt-oss-20b (high)openai/gpt-oss-20b:high | OpenAI | 43.5 | $0.03 / $0.13 | 1 of 6 benchmarks | 53.9% | |||||
| 256 | MiniMax M1minimax/minimax-m1 | MiniMax | 43.2 | $0.55 / $2.20 | 1 of 6 benchmarks | 1371 | |||||
| 257 | Qwen3.5-Flashqwen/qwen3.5-flash-02-23 | Qwen | 43.1 | $0.065 / $0.26 | 4 of 6 benchmarks | 10.0% | 0.0% | 84.4% | 1403 | ||
| 258 | Claude 3.7 Sonnet (16K)anthropic/claude-3-7-sonnet:16k | Anthropic | 43.1 | — | 3 of 6 benchmarks | 4.1% | 46.7% | 86.3% | |||
| 259 | Claude Sonnet 4 (16K)anthropic/claude-sonnet-4:16k | Anthropic | 43.0 | $3.00 / $15.00 | 1 of 6 benchmarks | 53.3% | |||||
| 260 | GPT-5.6 Terra (none)openai/gpt-5.6-terra:none | OpenAI | 43.0 | $1.00 / $6.00 | 1 of 6 benchmarks | 53.3% | |||||
| 261 | o1 (low)openai/o1:low | OpenAI | 43.0 | $15.00 / $60.00 | 1 of 6 benchmarks | 53.3% | |||||
| 262 | GLM 4.6z-ai/glm-4.6 | Z.ai | 43.0 | $0.50 / $2.00 | 3 of 6 benchmarks | 3.8% | 2.1% | 1420 | |||
| 263 | Grok 3 Mini Betax-ai/grok-3-mini-beta | xAI | 42.9 | — | 1 of 6 benchmarks | 1368 | |||||
| 264 | o3 Mini Highopenai/o3-mini-high | OpenAI | 42.9 | $1.10 / $4.40 | 6 of 6 benchmarks | 12.4% | 18.6% | 0.0% | 76.9% | 96.5% | 1405 |
| 265 | GLM 4.7 Flashz-ai/glm-4.7-flash | Z.ai | 42.8 | $0.06 / $0.40 | 1 of 6 benchmarks | 1366 | |||||
| 266 | Grok 3 Mini Beta (high)x-ai/grok-3-mini-beta:high | xAI | 42.7 | — | 5 of 6 benchmarks | 5.9% | 0.0% | 77.8% | 88.1% | 1387 | |
| 267 | Gemini 2.5 Flash Lite (thinking)google/gemini-2.5-flash-lite:thinking | 42.6 | $0.10 / $0.40 | 1 of 6 benchmarks | 1365 | ||||||
| 268 | Gemini 3.5 Flash Lite (minimal)google/gemini-3.5-flash-lite:minimal | 42.5 | $0.30 / $2.50 | 1 of 6 benchmarks | 51.1% | ||||||
| 269 | QwQ 32Bqwen/qwq-32b | Qwen | 42.5 | — | 1 of 6 benchmarks | 1364 | |||||
| 270 | GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | 42.4 | $0.40 / $1.60 | 4 of 6 benchmarks | 10.0% | 44.7% | 87.3% | 1354 | ||
| 271 | Gemini 2.5 Flash Lite (nothinking)google/gemini-2.5-flash-lite:nothinking | 42.4 | $0.10 / $0.40 | 1 of 6 benchmarks | 1364 | ||||||
| 272 | Claude Haiku 4.5anthropic/claude-haiku-4.5 | Anthropic | 42.3 | $1.00 / $5.00 | 4 of 6 benchmarks | 4.1% | 35.8% | 86.9% | 1398 | ||
| 273 | GLM 4.7z-ai/glm-4.7 | Z.ai | 42.3 | $0.40 / $1.75 | 4 of 6 benchmarks | 2.4% | 0.0% | 83.3% | 1428 | ||
| 274 | O1 Mini (high)openai/o1-mini:high | OpenAI | 42.2 | — | 3 of 6 benchmarks | 1.4% | 46.9% | 89.2% | |||
| 275 | Qwen 2.5 (max)qwen/qwen-2.5:max | Qwen | 42.2 | — | 1 of 6 benchmarks | 1363 | |||||
| 276 | Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | MoonshotAI | 42.2 | $0.60 / $2.50 | 2 of 6 benchmarks | 21.4% | 0.0% | ||||
| 277 | Step 3stepfun/step-3 | StepFun | 41.9 | — | 1 of 6 benchmarks | 1362 | |||||
| 278 | O1 Miniopenai/o1-mini | OpenAI | 41.8 | — | 1 of 6 benchmarks | 1362 | |||||
| 279 | Trinity Large Previewarcee-ai/trinity-large-preview | Arcee AI | 41.6 | — | 1 of 6 benchmarks | 1362 | |||||
| 280 | DeepSeek V4 Pro (none)deepseek/deepseek-v4-pro:none | DeepSeek | 41.6 | $1.168 / $2.336 | 1 of 6 benchmarks | 46.7% | |||||
| 281 | GPT-5 Nano (low)openai/gpt-5-nano:low | OpenAI | 41.6 | $0.05 / $0.40 | 1 of 6 benchmarks | 46.7% | |||||
| 282 | GPT-5 (minimal)openai/gpt-5:minimal | OpenAI | 41.6 | $1.25 / $10.00 | 1 of 6 benchmarks | 46.7% | |||||
| 283 | GPT-5.4 Nano (none)openai/gpt-5.4-nano:none | OpenAI | 41.6 | $0.20 / $1.25 | 1 of 6 benchmarks | 46.7% | |||||
| 284 | Claude Opus 4 (27K)anthropic/claude-opus-4:27k | Anthropic | 41.5 | $15.00 / $75.00 | 3 of 6 benchmarks | 4.1% | 4.2% | 64.4% | |||
| 285 | GLM 4.5Vz-ai/glm-4.5v | Z.ai | 41.5 | $0.60 / $1.80 | 1 of 6 benchmarks | 1360 | |||||
| 286 | MiniMax M2minimax/minimax-m2 | MiniMax | 41.2 | $0.255 / $1.02 | 1 of 6 benchmarks | 1354 | |||||
| 287 | Qwen3 30B A3Bqwen/qwen3-30b-a3b | Qwen | 40.9 | $0.12 / $0.50 | 1 of 6 benchmarks | 1352 | |||||
| 288 | Ling Flash 2.0inclusionai/ling-flash-2.0 | inclusionAI | 40.8 | — | 1 of 6 benchmarks | 1352 | |||||
| 289 | Gemini 3.1 Flash Lite (low)google/gemini-3.1-flash-lite:low | 40.7 | $0.25 / $1.50 | 1 of 6 benchmarks | 44.4% | ||||||
| 290 | o3 Mini (low)openai/o3-mini:low | OpenAI | 40.7 | $1.10 / $4.40 | 1 of 6 benchmarks | 44.4% | |||||
| 291 | Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | NVIDIA | 40.5 | $0.05 / $0.20 | 1 of 6 benchmarks | 1351 | |||||
| 292 | Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | 40.5 | $3.00 / $15.00 | 4 of 6 benchmarks | 9.3% | 2.1% | 35.6% | 1427 | ||
| 293 | Hunyuan TurboStencent/hunyuan-turbos | Tencent | 40.2 | — | 1 of 6 benchmarks | 1347 | |||||
| 294 | GPT-5.6 Luna (none)openai/gpt-5.6-luna:none | OpenAI | 40.2 | $0.10 / $0.60 | 1 of 6 benchmarks | 40.0% | |||||
| 295 | Ring Flash 2.0inclusionai/ring-flash-2.0 | inclusionAI | 39.9 | — | 1 of 6 benchmarks | 1340 | |||||
| 296 | Claude 3.7 Sonnet (64K)anthropic/claude-3-7-sonnet:64k | Anthropic | 39.8 | — | 4 of 6 benchmarks | 3.1% | 0.0% | 57.8% | 91.2% | ||
| 297 | O1 Mini (medium)openai/o1-mini:medium | OpenAI | 39.7 | — | 3 of 6 benchmarks | 1.7% | 44.7% | 84.3% | |||
| 298 | Mistral Small 2506mistralai/mistral-small-2506 | Mistral AI | 39.6 | — | 1 of 6 benchmarks | 1338 | |||||
| 299 | Gemini 3.1 Flash Lite (minimal)google/gemini-3.1-flash-lite:minimal | 39.5 | $0.25 / $1.50 | 1 of 6 benchmarks | 37.8% | ||||||
| 300 | gpt-oss-20bopenai/gpt-oss-20b | OpenAI | 39.5 | $0.03 / $0.13 | 1 of 6 benchmarks | 1336 | |||||
| 301 | Nova 2 Liteamazon/nova-2-lite-v1 | Amazon | 39.4 | $0.30 / $2.50 | 1 of 6 benchmarks | 1333 | |||||
| 302 | Gemini 2.0 Flash Lite Preview 02 05google/gemini-2.0-flash-lite-preview-02-05 | 39.2 | — | 1 of 6 benchmarks | 1326 | ||||||
| 303 | Granite 4.1 8Bibm-granite/granite-4.1-8b | IBM | 38.9 | $0.05 / $0.10 | 1 of 6 benchmarks | 1318 | |||||
| 304 | GPT-5 Nano (minimal)openai/gpt-5-nano:minimal | OpenAI | 38.9 | $0.05 / $0.40 | 1 of 6 benchmarks | 35.6% | |||||
| 305 | Gemma 3 12Bgoogle/gemma-3-12b-it | 38.8 | $0.05 / $0.15 | 1 of 6 benchmarks | 1318 | ||||||
| 306 | Claude Opus 4anthropic/claude-opus-4 | Anthropic | 38.5 | $15.00 / $75.00 | 5 of 6 benchmarks | 4.5% | 0.0% | 42.2% | 85.0% | 1404 | |
| 307 | OLMo 3 32B Thinkallenai/olmo-3-32b-think | Allen Institute for AI | 38.4 | — | 1 of 6 benchmarks | 1311 | |||||
| 308 | Command A (03-2025)cohere/command-a-03-2025 | Cohere | 38.1 | — | 1 of 6 benchmarks | 1309 | |||||
| 309 | Grok 3 Betax-ai/grok-3-beta | xAI | 38.0 | — | 5 of 6 benchmarks | 3.8% | 0.0% | 55.6% | 88.8% | 1373 | |
| 310 | GLM 5.2 (none)z-ai/glm-5.2:none | Z.ai | 37.8 | $0.63 / $1.98 | 1 of 6 benchmarks | 28.9% | |||||
| 311 | Claude Sonnet 4 (59K)anthropic/claude-sonnet-4:59k | Anthropic | 37.8 | $3.00 / $15.00 | 2 of 6 benchmarks | 0.0% | 68.9% | ||||
| 312 | GPT-5.4 Mini (none)openai/gpt-5.4-mini:none | OpenAI | 37.5 | $0.75 / $4.50 | 1 of 6 benchmarks | 26.7% | |||||
| 313 | OLMo 3.1 32B Instructallenai/olmo-3.1-32b-instruct | Allen Institute for AI | 37.5 | — | 1 of 6 benchmarks | 1304 | |||||
| 314 | Step 1o Turbo 202506stepfun/step-1o-turbo-202506 | StepFun | 37.1 | — | 1 of 6 benchmarks | 1297 | |||||
| 315 | Qwen3 235B A22Bqwen/qwen3-235b-a22b | Qwen | 36.7 | $0.455 / $1.82 | 3 of 6 benchmarks | 0.0% | 68.9% | 1392 | |||
| 316 | OLMo 3.1 32B Thinkallenai/olmo-3.1-32b-think | Allen Institute for AI | 36.6 | — | 1 of 6 benchmarks | 1296 | |||||
| 317 | Magistral Medium 2506mistralai/magistral-medium-2506 | Mistral AI | 35.9 | — | 1 of 6 benchmarks | 1285 | |||||
| 318 | GPT-4.1openai/gpt-4.1 | OpenAI | 35.9 | $2.00 / $8.00 | 5 of 6 benchmarks | 5.5% | 0.0% | 38.3% | 83.0% | 1374 | |
| 319 | Gemma 3 27Bgoogle/gemma-3-27b-it | 35.8 | $0.08 / $0.45 | 3 of 6 benchmarks | 22.2% | 74.0% | 1323 | ||||
| 320 | Gemini 1.5 Pro 002google/gemini-1.5-pro-002 | 35.7 | — | 3 of 6 benchmarks | 23.1% | 70.4% | 1339 | ||||
| 321 | Claude Sonnet 4anthropic/claude-sonnet-4 | Anthropic | 35.7 | $3.00 / $15.00 | 5 of 6 benchmarks | 4.1% | 0.0% | 28.9% | 84.4% | 1389 | |
| 322 | Gemini 2.0 Flash 001google/gemini-2.0-flash-001 | 35.7 | — | 4 of 6 benchmarks | 1.7% | 31.1% | 82.2% | 1356 | |||
| 323 | Hunyuan Large Visiontencent/hunyuan-large-vision | Tencent | 35.6 | — | 1 of 6 benchmarks | 1280 | |||||
| 324 | Claude Opus 4.1 (16K)anthropic/claude-opus-4.1:16k | Anthropic | 35.1 | $15.00 / $75.00 | 2 of 6 benchmarks | 0.0% | 64.4% | ||||
| 325 | Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct | Qwen | 35.1 | — | 1 of 6 benchmarks | 1270 | |||||
| 326 | Amazon Nova Pro v1.0amazon/amazon-nova-pro-v1.0 | Amazon | 34.8 | — | 1 of 6 benchmarks | 1269 | |||||
| 327 | GPT-5.1 (none)openai/gpt-5.1:none | OpenAI | 34.6 | $1.25 / $10.00 | 2 of 6 benchmarks | 2.1% | 37.8% | ||||
| 328 | Gemma 3n E4Bgoogle/gemma-3n-e4b-it | 34.5 | — | 1 of 6 benchmarks | 1260 | ||||||
| 329 | Gemma 3 4Bgoogle/gemma-3-4b-it | 34.2 | — | 1 of 6 benchmarks | 1254 | ||||||
| 330 | Amazon Nova Lite v1.0amazon/amazon-nova-lite-v1.0 | Amazon | 33.9 | — | 1 of 6 benchmarks | 1244 | |||||
| 331 | Claude 3.7 Sonnetanthropic/claude-3-7-sonnet | Anthropic | 33.9 | — | 4 of 6 benchmarks | 3.1% | 21.9% | 68.2% | 1363 | ||
| 332 | DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | DeepSeek | 33.7 | $0.27 / $1.12 | 4 of 6 benchmarks | 0.0% | 37.8% | 75.5% | 1369 | ||
| 333 | Command R+ (08-2024)cohere/command-r-plus-08-2024 | Cohere | 33.6 | $2.50 / $10.00 | 1 of 6 benchmarks | 1231 | |||||
| 334 | Mistral Medium 2505mistralai/mistral-medium-2505 | Mistral AI | 33.6 | — | 4 of 6 benchmarks | 0.3% | 32.2% | 81.6% | 1348 | ||
| 335 | DeepSeek V3deepseek/deepseek-chat | DeepSeek | 33.3 | $0.2574 / $1.0287 | 4 of 6 benchmarks | 1.7% | 48.9% | 64.8% | 1311 | ||
| 336 | Claude Opus 4.1 (32K)anthropic/claude-opus-4.1:32k | Anthropic | 33.3 | $15.00 / $75.00 | 2 of 6 benchmarks | 12.6% | 2.4% | ||||
| 337 | OLMo 2 0325 32B Instructallenai/olmo-2-0325-32b-instruct | Allen Institute for AI | 33.2 | — | 1 of 6 benchmarks | 1227 | |||||
| 338 | Amazon Nova Micro v1.0amazon/amazon-nova-micro-v1.0 | Amazon | 33.1 | — | 1 of 6 benchmarks | 1224 | |||||
| 339 | QwQ 32B Previewqwen/qwq-32b-preview | Qwen | 32.9 | — | 1 of 6 benchmarks | 1210 | |||||
| 340 | Command R (08-2024)cohere/command-r-08-2024 | Cohere | 32.8 | $0.15 / $0.60 | 1 of 6 benchmarks | 1207 | |||||
| 341 | GLM 4.5z-ai/glm-4.5 | Z.ai | 32.7 | $0.60 / $2.20 | 3 of 6 benchmarks | 0.0% | 0.0% | 1413 | |||
| 342 | Qwen (max)qwen/qwen:max | Qwen | 32.6 | — | 3 of 6 benchmarks | 1.0% | 16.1% | 67.2% | |||
| 343 | Mistral Nemomistralai/mistral-nemo | Mistral | 32.6 | $0.019 / $0.03 | 1 of 6 benchmarks | 10.8% | |||||
| 344 | Grok 2 1212x-ai/grok-2-1212 | xAI | 31.5 | — | 3 of 6 benchmarks | 0.7% | 11.5% | 63.5% | |||
| 345 | GPT 4 1106 Previewopenai/gpt-4-1106-preview | OpenAI | 31.2 | — | 2 of 6 benchmarks | 40.0% | 1303 | ||||
| 346 | Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct | Meta | 31.1 | — | 4 of 6 benchmarks | 0.7% | 20.6% | 73.0% | 1317 | ||
| 347 | Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct | Qwen | 31.1 | $0.36 / $0.40 | 3 of 6 benchmarks | 8.1% | 63.2% | 1296 | |||
| 348 | GPT-4.1 Nanoopenai/gpt-4.1-nano | OpenAI | 29.8 | $0.10 / $0.40 | 4 of 6 benchmarks | 1.0% | 28.9% | 70.0% | 1274 | ||
| 349 | Magistral Small 2506mistralai/magistral-small-2506 | Mistral AI | 29.3 | — | 2 of 6 benchmarks | 0.0% | 30.0% | ||||
| 350 | GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | OpenAI | 29.1 | $5.00 / $15.00 | 3 of 6 benchmarks | 6.3% | 51.0% | 1305 | |||
| 351 | GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | OpenAI | 28.5 | $0.15 / $0.60 | 3 of 6 benchmarks | 6.9% | 52.6% | 1276 | |||
| 352 | GPT-4 Turboopenai/gpt-4-turbo | OpenAI | 27.8 | $10.00 / $30.00 | 3 of 6 benchmarks | 6.7% | 46.7% | 1296 | |||
| 353 | Mistral Large 2407mistralai/mistral-large-2407 | Mistral | 27.3 | $2.00 / $6.00 | 3 of 6 benchmarks | 8.5% | 44.8% | 1288 | |||
| 354 | GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | OpenAI | 27.2 | $2.50 / $10.00 | 3 of 6 benchmarks | 0.3% | 6.3% | 49.8% | |||
| 355 | Mistral Small 3.1 24B Instruct 2503mistralai/mistral-small-3.1-24b-instruct-2503 | Mistral AI | 26.9 | — | 3 of 6 benchmarks | 5.8% | 46.8% | 1278 | |||
| 356 | Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct | Mistral | 26.8 | $2.00 / $6.00 | 2 of 6 benchmarks | 24.2% | 1228 | ||||
| 357 | Claude 3.5 Sonnetanthropic/claude-3-5-sonnet | Anthropic | 26.7 | — | 5 of 6 benchmarks | 1.0% | 0.0% | 8.5% | 57.0% | 1351 | |
| 358 | Gemini 1.5 Pro 001google/gemini-1.5-pro-001 | 26.7 | — | 3 of 6 benchmarks | 6.8% | 40.8% | 1299 | ||||
| 359 | Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct | Meta | 26.5 | — | 4 of 6 benchmarks | 0.0% | 7.8% | 62.3% | 1308 | ||
| 360 | GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | OpenAI | 26.4 | $2.50 / $10.00 | 4 of 6 benchmarks | 0.3% | 6.4% | 53.3% | 1309 | ||
| 361 | Claude 3 Opusanthropic/claude-3-opus | Anthropic | 26.1 | — | 3 of 6 benchmarks | 4.7% | 37.5% | 1312 | |||
| 362 | Gemini 1.5 Flash 002google/gemini-1.5-flash-002 | 26.1 | — | 4 of 6 benchmarks | 0.0% | 16.3% | 61.9% | 1289 | |||
| 363 | Gemini 1.5 Flash 8B 001google/gemini-1.5-flash-8b-001 | 26.0 | — | 2 of 6 benchmarks | 4.6% | 1230 | |||||
| 364 | Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | Meta | 25.9 | $0.10 / $0.32 | 3 of 6 benchmarks | 5.1% | 41.6% | 1296 | |||
| 365 | Mistral Small 24B Instruct 2501mistralai/mistral-small-24b-instruct-2501 | Mistral AI | 25.3 | — | 3 of 6 benchmarks | 5.3% | 44.8% | 1262 | |||
| 366 | Claude 2anthropic/claude-2 | Anthropic | 25.1 | — | 2 of 6 benchmarks | 2.5% | 11.7% | ||||
| 367 | Mistral Large 2411mistralai/mistral-large-2411 | Mistral AI | 24.9 | — | 4 of 6 benchmarks | 0.3% | 7.8% | 50.3% | 1282 | ||
| 368 | Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | Meta | 23.3 | $0.40 / $0.40 | 3 of 6 benchmarks | 3.6% | 36.7% | 1269 | |||
| 369 | Claude 3.5 Haikuanthropic/claude-3-5-haiku | Anthropic | 23.1 | — | 4 of 6 benchmarks | 0.3% | 4.3% | 46.4% | 1286 | ||
| 370 | Gemini 1.5 Flash 001google/gemini-1.5-flash-001 | 22.8 | — | 3 of 6 benchmarks | 3.9% | 25.1% | 1258 | ||||
| 371 | Claude 3 Sonnetanthropic/claude-3-sonnet | Anthropic | 21.5 | — | 3 of 6 benchmarks | 2.5% | 18.2% | 1253 | |||
| 372 | Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | Meta | 20.9 | $0.05 / $0.08 | 3 of 6 benchmarks | 2.5% | 22.9% | 1189 | |||
| 373 | Claude 3 Haikuanthropic/claude-3-haiku | Anthropic | 20.8 | $0.25 / $1.25 | 3 of 6 benchmarks | 1.8% | 14.9% | 1231 |
How this ranks
Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.
A model scored on fewer than 3 of the 6 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.
Data sources
Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.
- Epoch AI Benchmarking HubCC BY 4.0
Benchmark runs by Epoch AI, from the AI Benchmarking Hub.
- LMArenaCC BY 4.0
Arena ratings by LMArena, from the public leaderboard dataset.