Best LLM for agents
Agentic tool use covers whether a model drives tools to a finished outcome, stays steerable, and recovers when a command fails.
Rumeqo runs these models inside your team rooms. See what each one costs.
| rank | model | vendor | composite | pricein / out | benchmarks | % | score | score | score | score | score |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 (high)anthropic/claude-fable-5:high | Anthropic | 81.7 | $10.00 / $50.00 | 5 of 6 benchmarks | 12.0 | 9.2 | 10.8 | 1.2 | 13.7 | |
| 2 | Claude Opus 5 (high)anthropic/claude-opus-5:high | Anthropic | 77.6 | $5.00 / $25.00 | 5 of 6 benchmarks | 12.0 | 10.4 | 15.2 | 1.1 | 13.6 | |
| 3 | GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh | OpenAI | 76.4 | $5.00 / $30.00 | 5 of 6 benchmarks | 10.7 | 8.9 | 9.8 | 1.2 | 10.5 | |
| 4 | Claude Opus 5 (max)anthropic/claude-opus-5:max | Anthropic | 76.1 | $5.00 / $25.00 | 5 of 6 benchmarks | 11.9 | 6.7 | 18.1 | 1.1 | 14.1 | |
| 5 | Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | 74.5 | $5.00 / $25.00 | 6 of 6 benchmarks | 75.6% | 6.7 | 8.2 | 5.0 | 1.2 | 11.1 |
| 6 | GPT-5.5 (xhigh)openai/gpt-5.5:xhigh | OpenAI | 74.4 | $5.00 / $30.00 | 5 of 6 benchmarks | 8.7 | 8.9 | 4.9 | 1.2 | 14.4 | |
| 7 | Kimi K3 (max)moonshotai/kimi-k3:max | MoonshotAI | 73.6 | $3.00 / $15.00 | 5 of 6 benchmarks | 10.4 | 7.5 | 15.3 | 1.2 | 7.9 | |
| 8 | GPT-5.5 (high)openai/gpt-5.5:high | OpenAI | 72.8 | $5.00 / $30.00 | 5 of 6 benchmarks | 7.6 | 8.7 | 3.9 | 1.2 | 12.9 | |
| 9 | Claude Opus 4.7 (thinking)anthropic/claude-opus-4.7:thinking | Anthropic | 69.6 | $5.00 / $25.00 | 5 of 6 benchmarks | 8.2 | 7.9 | 6.7 | 1.1 | 12.7 | |
| 10 | Claude Opus 4.7anthropic/claude-opus-4.7 | Anthropic | 69.1 | $5.00 / $25.00 | 5 of 6 benchmarks | 7.6 | 10.3 | 5.5 | 1.1 | 9.9 | |
| 11 | GPT-5.5openai/gpt-5.5 | OpenAI | 68.7 | $5.00 / $30.00 | 5 of 6 benchmarks | 6.2 | 7.0 | 3.7 | 1.2 | 11.5 | |
| 12 | Grok 4.5x-ai/grok-4.5 | SpaceXAI | 68.7 | $2.00 / $6.00 | 5 of 6 benchmarks | 5.7 | 6.4 | 5.9 | 1.2 | 10.8 | |
| 13 | Claude Opus 4.5 (high)anthropic/claude-opus-4.5:high | Anthropic | 66.7 | $5.00 / $25.00 | 1 of 6 benchmarks | 76.8% | |||||
| 14 | GLM 5.2 (max)z-ai/glm-5.2:max | Z.ai | 66.5 | $0.63 / $1.98 | 5 of 6 benchmarks | 6.7 | 6.1 | 8.4 | 1.2 | 5.8 | |
| 15 | Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high | 65.6 | $0.50 / $3.00 | 1 of 6 benchmarks | 75.8% | ||||||
| 16 | MiniMax M2.5 (high)minimax/minimax-m2.5:high | MiniMax | 65.6 | $0.22 / $0.90 | 1 of 6 benchmarks | 75.8% | |||||
| 17 | GPT-5.4 (high)openai/gpt-5.4:high | OpenAI | 65.2 | $2.50 / $15.00 | 5 of 6 benchmarks | 5.0 | 6.2 | 4.7 | 1.2 | 9.2 | |
| 18 | Claude Opus 4.5 (medium)anthropic/claude-opus-4.5:medium | Anthropic | 63.7 | $5.00 / $25.00 | 1 of 6 benchmarks | 74.4% | |||||
| 19 | Gemini 3 Pro Previewgoogle/gemini-3-pro-preview | 63.0 | — | 1 of 6 benchmarks | 74.2% | ||||||
| 20 | Claude Opus 4.8 (thinking)anthropic/claude-opus-4.8:thinking | Anthropic | 61.8 | $5.00 / $25.00 | 5 of 6 benchmarks | 9.5 | 8.1 | 9.1 | -0.9 | 9.1 | |
| 21 | Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high | Anthropic | 61.8 | $2.00 / $10.00 | 5 of 6 benchmarks | 7.4 | 5.7 | 4.0 | 1.0 | 10.8 | |
| 22 | GPT-5.2-Codexopenai/gpt-5.2-codex | OpenAI | 61.5 | $1.75 / $14.00 | 1 of 6 benchmarks | 72.8% | |||||
| 23 | GPT-5.2 (high)openai/gpt-5.2:high | OpenAI | 61.5 | $1.75 / $14.00 | 1 of 6 benchmarks | 72.8% | |||||
| 24 | GLM 5 (high)z-ai/glm-5:high | Z.ai | 61.5 | $0.95 / $2.55 | 1 of 6 benchmarks | 72.8% | |||||
| 25 | GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh | OpenAI | 61.5 | $0.10 / $0.60 | 5 of 6 benchmarks | 4.3 | 1.7 | -0.9 | 1.2 | 11.3 | |
| 26 | DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high | DeepSeek | 61.0 | $0.14 / $0.28 | 5 of 6 benchmarks | 3.9 | 2.8 | 9.1 | 1.2 | 4.1 | |
| 27 | Claude Sonnet 4.5 (high)anthropic/claude-sonnet-4.5:high | Anthropic | 60.0 | $3.00 / $15.00 | 1 of 6 benchmarks | 71.4% | |||||
| 28 | Kimi K2.5 (high)moonshotai/kimi-k2.5:high | MoonshotAI | 59.3 | $0.57 / $2.85 | 1 of 6 benchmarks | 70.8% | |||||
| 29 | GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh | OpenAI | 59.0 | $1.00 / $6.00 | 5 of 6 benchmarks | 3.4 | 5.1 | -1.3 | 1.2 | 9.3 | |
| 30 | Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | 58.6 | $3.00 / $15.00 | 1 of 6 benchmarks | 70.6% | |||||
| 31 | DeepSeek V3.2 (high)deepseek/deepseek-v3.2:high | DeepSeek | 57.8 | $0.269 / $0.40 | 1 of 6 benchmarks | 70.0% | |||||
| 32 | Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | 57.8 | $5.00 / $25.00 | 5 of 6 benchmarks | 2.5 | 8.2 | 8.7 | -28.6 | 10.5 | |
| 33 | Gemini 3 Pro Preview (high)google/gemini-3-pro-preview:high | 57.1 | — | 1 of 6 benchmarks | 69.6% | ||||||
| 34 | GPT-5.2openai/gpt-5.2 | OpenAI | 56.3 | $1.75 / $14.00 | 1 of 6 benchmarks | 69.0% | |||||
| 35 | Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | 55.9 | $3.00 / $15.00 | 5 of 6 benchmarks | 3.0 | 2.2 | -0.2 | 1.1 | 10.9 | |
| 36 | Claude Opus 4anthropic/claude-opus-4 | Anthropic | 55.6 | $15.00 / $75.00 | 1 of 6 benchmarks | 67.6% | |||||
| 37 | Muse Spark 1.1meta/muse-spark-1.1 | Meta | 55.3 | $1.25 / $4.25 | 5 of 6 benchmarks | 1.1 | -3.2 | 6.7 | 1.1 | 5.7 | |
| 38 | Claude Haiku 4.5 (high)anthropic/claude-haiku-4.5:high | Anthropic | 54.8 | $1.00 / $5.00 | 1 of 6 benchmarks | 66.6% | |||||
| 39 | GPT-5.1-Codex (medium)openai/gpt-5.1-codex:medium | OpenAI | 53.7 | $1.25 / $10.00 | 1 of 6 benchmarks | 66.0% | |||||
| 40 | GPT-5.1 (medium)openai/gpt-5.1:medium | OpenAI | 53.7 | $1.25 / $10.00 | 1 of 6 benchmarks | 66.0% | |||||
| 41 | Kimi K2.7 Codemoonshotai/kimi-k2.7-code | MoonshotAI | 53.1 | $0.67 / $3.40 | 5 of 6 benchmarks | 1.0 | -2.0 | 4.6 | 1.2 | -1.6 | |
| 42 | GPT-5 (medium)openai/gpt-5:medium | OpenAI | 52.6 | $1.25 / $10.00 | 1 of 6 benchmarks | 65.0% | |||||
| 43 | Claude Sonnet 4anthropic/claude-sonnet-4 | Anthropic | 51.9 | $3.00 / $15.00 | 1 of 6 benchmarks | 64.9% | |||||
| 44 | Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | MoonshotAI | 51.1 | $0.60 / $2.50 | 1 of 6 benchmarks | 63.4% | |||||
| 45 | MiniMax M2minimax/minimax-m2 | MiniMax | 50.4 | $0.255 / $1.02 | 1 of 6 benchmarks | 61.0% | |||||
| 46 | DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking | DeepSeek | 49.7 | $0.269 / $0.40 | 1 of 6 benchmarks | 60.0% | |||||
| 47 | Kimi K2.6moonshotai/kimi-k2.6 | MoonshotAI | 49.1 | $0.95 / $4.00 | 5 of 6 benchmarks | -0.6 | -1.0 | 0.7 | 1.2 | -5.8 | |
| 48 | GPT-5 Mini (medium)openai/gpt-5-mini:medium | OpenAI | 48.9 | $0.25 / $2.00 | 1 of 6 benchmarks | 59.8% | |||||
| 49 | o3openai/o3 | OpenAI | 48.2 | $2.00 / $8.00 | 1 of 6 benchmarks | 58.4% | |||||
| 50 | Devstral Small 2512mistralai/devstral-small-2512 | Mistral AI | 47.4 | — | 1 of 6 benchmarks | 56.4% | |||||
| 51 | GPT-5 Miniopenai/gpt-5-mini | OpenAI | 46.7 | $0.25 / $2.00 | 1 of 6 benchmarks | 56.2% | |||||
| 52 | Qwen3.7 Maxqwen/qwen3.7-max | Qwen | 46.6 | $1.475 / $4.425 | 5 of 6 benchmarks | -0.0 | -0.2 | -0.1 | 0.5 | 4.7 | |
| 53 | Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct | Qwen | 45.6 | — | 1 of 6 benchmarks | 55.4% | |||||
| 54 | GLM 4.6z-ai/glm-4.6 | Z.ai | 45.6 | $0.50 / $2.00 | 1 of 6 benchmarks | 55.4% | |||||
| 55 | Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 45.5 | $2.00 / $12.00 | 5 of 6 benchmarks | -0.6 | 3.1 | 1.4 | 0.9 | -11.6 | ||
| 56 | GLM 4.5z-ai/glm-4.5 | Z.ai | 44.5 | $0.60 / $2.20 | 1 of 6 benchmarks | 54.2% | |||||
| 57 | Devstral 2mistralai/devstral-2 | Mistral AI | 43.7 | — | 1 of 6 benchmarks | 53.8% | |||||
| 58 | GLM 5.1z-ai/glm-5.1 | Z.ai | 43.5 | $1.40 / $4.40 | 5 of 6 benchmarks | 0.5 | 1.5 | 1.7 | -0.5 | -1.0 | |
| 59 | Gemini 2.5 Progoogle/gemini-2.5-pro | 43.0 | $1.25 / $10.00 | 1 of 6 benchmarks | 53.6% | ||||||
| 60 | DeepSeek V4 Prodeepseek/deepseek-v4-pro | DeepSeek | 42.7 | $1.168 / $2.336 | 5 of 6 benchmarks | -0.1 | 0.6 | -2.2 | 0.3 | 4.0 | |
| 61 | Claude 3.7 Sonnetanthropic/claude-3-7-sonnet | Anthropic | 42.3 | — | 1 of 6 benchmarks | 52.8% | |||||
| 62 | Gemini 3.5 Flash (high)google/gemini-3.5-flash:high | 41.6 | $1.50 / $9.00 | 5 of 6 benchmarks | -0.4 | -0.7 | 0.5 | 0.3 | -1.9 | ||
| 63 | o4 Miniopenai/o4-mini | OpenAI | 41.5 | $1.10 / $4.40 | 1 of 6 benchmarks | 45.0% | |||||
| 64 | Kimi K2 Instructmoonshotai/kimi-k2-instruct | Moonshot AI | 40.8 | — | 1 of 6 benchmarks | 43.8% | |||||
| 65 | Gemini 3.6 Flashgoogle/gemini-3.6-flash | 40.1 | $1.50 / $7.50 | 5 of 6 benchmarks | -2.8 | -5.0 | -0.8 | 1.1 | -3.6 | ||
| 66 | GPT-4.1openai/gpt-4.1 | OpenAI | 40.0 | $2.00 / $8.00 | 1 of 6 benchmarks | 39.6% | |||||
| 67 | Qwen3.7 Plusqwen/qwen3.7-plus | Qwen | 39.3 | $0.32 / $1.28 | 5 of 6 benchmarks | -1.8 | -4.9 | -1.0 | 0.2 | 5.9 | |
| 68 | GPT-5 Nano (medium)openai/gpt-5-nano:medium | OpenAI | 39.3 | $0.05 / $0.40 | 1 of 6 benchmarks | 34.8% | |||||
| 69 | Gemini 2.5 Flashgoogle/gemini-2.5-flash | 38.6 | $0.30 / $2.50 | 1 of 6 benchmarks | 28.7% | ||||||
| 70 | MiniMax M3minimax/minimax-m3 | MiniMax | 38.2 | $0.30 / $1.20 | 5 of 6 benchmarks | -2.5 | -4.8 | -5.8 | 0.6 | 5.5 | |
| 71 | gpt-oss-120bopenai/gpt-oss-120b | OpenAI | 37.8 | $0.03 / $0.17 | 1 of 6 benchmarks | 26.0% | |||||
| 72 | GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | 37.1 | $0.40 / $1.60 | 1 of 6 benchmarks | 23.9% | |||||
| 73 | GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | OpenAI | 36.3 | $2.50 / $10.00 | 1 of 6 benchmarks | 21.6% | |||||
| 74 | MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | Xiaomi | 35.7 | $0.435 / $0.87 | 5 of 6 benchmarks | -2.2 | -2.4 | -3.2 | -0.1 | 1.7 | |
| 75 | Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct | Meta | 35.6 | — | 1 of 6 benchmarks | 21.0% | |||||
| 76 | DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | DeepSeek | 35.6 | $0.14 / $0.28 | 5 of 6 benchmarks | -2.3 | -1.2 | -2.2 | -1.3 | 2.5 | |
| 77 | Gemini 2.0 Flash 001google/gemini-2.0-flash-001 | 34.8 | — | 1 of 6 benchmarks | 13.5% | ||||||
| 78 | Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct | Meta | 34.1 | — | 1 of 6 benchmarks | 9.1% | |||||
| 79 | Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium | 33.6 | $1.50 / $9.00 | 5 of 6 benchmarks | -3.6 | -3.9 | -8.2 | 0.4 | -0.6 | ||
| 80 | Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct | Qwen | 33.4 | — | 1 of 6 benchmarks | 9.0% | |||||
| 81 | Hy3tencent/hy3 | Tencent | 33.2 | $0.132 / $0.528 | 5 of 6 benchmarks | -1.3 | -8.5 | -2.4 | -1.3 | 3.5 | |
| 82 | Inklingthinkingmachines/inkling | Thinking Machines | 32.3 | $0.95 / $4.05 | 5 of 6 benchmarks | -6.7 | -11.8 | -12.5 | 0.5 | 6.4 | |
| 83 | Grok 4.3 (high)x-ai/grok-4.3:high | SpaceXAI | 30.4 | $1.25 / $2.50 | 5 of 6 benchmarks | -8.5 | -7.0 | -9.6 | 1.0 | -13.0 | |
| 84 | Grok Build 0.1x-ai/grok-build-0.1 | SpaceXAI | 29.0 | $1.00 / $2.00 | 5 of 6 benchmarks | -9.0 | -8.8 | -5.6 | 0.9 | -20.7 | |
| 85 | Grok 4.3x-ai/grok-4.3 | SpaceXAI | 27.7 | $1.25 / $2.50 | 5 of 6 benchmarks | -14.6 | -6.0 | -10.9 | 1.1 | -41.1 | |
| 86 | Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 27.5 | $0.50 / $3.00 | 5 of 6 benchmarks | -8.6 | -3.9 | -7.5 | 0.2 | -21.1 | ||
| 87 | MiniMax M2.7minimax/minimax-m2.7 | MiniMax | 25.2 | $0.30 / $1.20 | 5 of 6 benchmarks | -11.1 | -13.5 | -11.3 | 1.0 | -16.7 | |
| 88 | Mistral Medium 3.5mistralai/mistral-medium-3-5 | Mistral | 25.2 | $1.50 / $7.50 | 5 of 6 benchmarks | -6.9 | -11.2 | -9.4 | -3.0 | -1.2 | |
| 89 | Solar Pro 4upstage/solar-pro4 | Upstage | 23.6 | $0.03 / $0.12 | 5 of 6 benchmarks | -12.1 | -13.6 | -7.3 | 0.3 | -16.5 | |
| 90 | Gemma 4 31Bgoogle/gemma-4-31b-it | 23.0 | $0.10 / $0.34 | 5 of 6 benchmarks | -18.2 | -9.3 | 0.1 | -30.9 | -48.6 | ||
| 91 | Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 22.4 | $0.30 / $2.50 | 5 of 6 benchmarks | -10.2 | -9.9 | -13.9 | -0.4 | -13.0 | ||
| 92 | Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | NVIDIA | 20.2 | $0.60 / $3.60 | 5 of 6 benchmarks | -14.6 | -19.7 | -15.8 | 0.5 | -23.7 |
How this ranks
Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.
A model scored on fewer than 3 of the 6 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.
Data sources
Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.