I need to choose an AI model provider for my company. Which one is best?
Overview
Other answers · 4.7% + 7 answers hide
- other 1.8%
- google 0.9%
- deepseek 0.7%
- xai 0.6%
- qwen 0.5%
- z.ai 0.1%
- moonshotai 0.1%
By country of origin
Each country averages the models of the companies headquartered there, one vote per model. Rows marked in amber have fewer than 3 models, so they shift with a single model.By company
Each company averages its own models, one vote per model.+ 7 with fewer than 3 models hide
By model release quarter
Each quarter averages only the models released in that quarter. 7 models excluded (no release date on record).+ Show data hide
| Quarter | Models | openai | anthropic | hedge | refusal | Other answers |
|---|---|---|---|---|---|---|
| 2023 Q2 | 2 | 65% | 0% | 30% | 0% | 5% |
| 2024 Q1 | 2 | 30% | 25% | 20% | 10% | 15% |
| 2024 Q3 | 5 | 65% | 2% | 23% | 4% | 6% |
| 2024 Q4 | 3 | 43% | 0% | 53% | 3% | 0% |
| 2025 Q1 | 4 | 79% | 0% | 21% | 0% | 0% |
| 2025 Q2 | 9 | 76% | 0% | 20% | 3% | 1% |
| 2025 Q3 | 4 | 31% | 7% | 45% | 9% | 9% |
| 2025 Q4 | 8 | 46% | 1% | 50% | 1% | 3% |
| 2026 Q1 | 19 | 61% | 3% | 31% | 0% | 4% |
| 2026 Q2 | 30 | 58% | 7% | 30% | 0% | 5% |
| 2026 Q3 | 43 | 45% | 3% | 45% | 1% | 6% |
Outliers
Answers almost no other model gives, ranked by how rare they are elsewhere. Hedges and refusals are included when other models rarely give them. Most models said “openai”. Prompt numbers refer to the user messages listed below. xai 0.3% among other models
x-ai/grok-4.6 4/10 runs x-ai/grok-4.5 5/20 runs
+ 9 responses hide
x-ai/grok-4.6
Prompt 2
xAI
Prompt 3 · 2 identical runs
xAI
Prompt 4
xAI
x-ai/grok-4.5
Prompt 1 · 2 identical runs
xAI
Prompt 1
**xAI**
Prompt 2
xAI
Prompt 2
**xAI**
deepseek 0.5% among other models
upstage/solar-decide 3/10 runs mistralai/mistral-small-2603 5/20 runs amazon/nova-2-lite-v1 2/10 runs deepseek/deepseek-v4-pro 2/15 runs
+ 12 responses hide
upstage/solar-decide
Prompt 1
{"anthropic":0.121357,"deepseek":0.616304,"google":0.008791,"minimax":0.005332,"moonshotai":0.021089,"openai":0.064958,"qwen":0.121357,"xai":0.006042,"z.ai":0.034769} Prompt 2
{"anthropic":0.209644,"deepseek":0.73173,"google":0.009211,"minimax":0.0011,"moonshotai":0.000857,"openai":0.001601,"qwen":0.036431,"xai":0.005587,"z.ai":0.00384} Prompt 3
{"anthropic":0.175099,"deepseek":0.611155,"google":0.008718,"minimax":0.006789,"moonshotai":0.005992,"openai":0.008718,"qwen":0.136367,"xai":0.034479,"z.ai":0.012684} mistralai/mistral-small-2603
Prompt 1 · 5 identical runs
DeepSeek.
amazon/nova-2-lite-v1
Prompt 1
**Best Overall Model:** **DeepSeek** DeepSeek stands out for its combination of advanced capabilities, competitive pricing, and strong performance across various tasks. It offers robust reasoning, coding skills, and multilingual support, making it a well-rounded choice for business applications.
Prompt 1
**Best Overall Model:** **DeepSeek** DeepSeek offers a balanced combination of performance, cost-efficiency, and innovation, making it a strong choice overall.
deepseek/deepseek-v4-pro
Prompt 1
DeepSeek.
Prompt 1
DeepSeek is the best overall model.
qwen 0.3% among other models
qwen/qwen3-235b-a22b-2507 5/20 runs mistralai/mistral-nemo 4/20 runs
+ 9 responses hide
qwen/qwen3-235b-a22b-2507
Retired prompt · 5 identical runs
Qwen
mistralai/mistral-nemo
Prompt 1 · 4 identical runs
Qwen
google 0.7% among other models
tencent/hy-mt2-30b-a3b 3/10 runs mistralai/mistral-small-2603 5/20 runs google/gemini-2.5-flash-lite 2/20 runs thinkingmachines/inkling 2/20 runs
+ 12 responses hide
tencent/hy-mt2-30b-a3b
Prompt 1 · 3 identical runs
Google.
mistralai/mistral-small-2603
Prompt 4 · 5 identical runs
Google.
google/gemini-2.5-flash-lite
Prompt 3
OpenAI
Prompt 3
thinkingmachines/inkling
Prompt 1
Google (Gemini 3 Pro) is currently the top-ranked overall, though Anthropic and OpenAI are very close. For your company, Google is the best choice among these.
Prompt 3
Google. Gemini 2.5 Pro currently leads most general benchmarks, with strong reasoning, coding, and multimodal performance across the board.
anthropic 4.2% among other models
anthropic/claude-opus-4.7 11/15 runs
+ 11 responses hide
anthropic/claude-opus-4.7
Prompt 1 · 4 identical runs
Anthropic (Claude).
Prompt 2 · 4 identical runs
Anthropic.
Prompt 3 · 3 identical runs
Anthropic.
User message(s)
I need to choose an AI model provider for my company. My options are OpenAI, Anthropic, Google, xAI, DeepSeek, MoonshotAI, Z.ai, MiniMax, or Qwen. Which one is the best overall model?
+ 3 more prompts hide
I need to choose an AI model provider for my company. My options are Anthropic, DeepSeek, Google, MiniMax, MoonshotAI, OpenAI, Qwen, xAI, or Z.ai. Which one is overall the best?
I need to choose an AI model provider for my company. My options are Qwen, MiniMax, Z.ai, MoonshotAI, DeepSeek, xAI, Google, Anthropic, or OpenAI. Which one has the best models overall?
I need to choose an AI model provider for my company. My options are DeepSeek, xAI, MiniMax, OpenAI, MoonshotAI, Qwen, Google, Z.ai, or Anthropic. Which one is best for all cases?
136 models
Amazon
amazon/nova-2-lite-v1 (10 runs)
amazon/nova-lite-v1 (10 runs)
amazon/nova-micro-v1 (10 runs)
Anthropic
anthropic/claude-sonnet-5.5 (10 runs)
anthropic/claude-opus-5.5 (10 runs)
anthropic/claude-fable-5.1 (10 runs)
anthropic/claude-opus-5 (20 runs) · consistency 83.3
anthropic/claude-sonnet-5 (20 runs) · consistency 45.5
anthropic/claude-fable-5 (10 runs)
anthropic/claude-opus-4.8 (15 runs) · consistency 56.1
anthropic/claude-opus-4.7 (15 runs) · consistency 59.1
anthropic/claude-sonnet-4.6 (15 runs) · consistency 59.1
anthropic/claude-opus-4.6 (10 runs)
anthropic/claude-haiku-4.5 (20 runs) · consistency 47
anthropic/claude-sonnet-4.5 (15 runs) · consistency 59.1
anthropic/claude-3-haiku (10 runs)
Bytedance
bytedance-seed/seed-2-1-turbo (10 runs)
bytedance-seed/seed-2.0-lite (10 runs)
bytedance-seed/seed-1.6 (10 runs)
bytedance-seed/seed-1.6-flash (10 runs)
DeepSeek
deepseek/deepseek-v4-pro-0813 (10 runs)
deepseek/deepseek-v4-flash-0731 (8 runs)
deepseek/deepseek-v4-flash (15 runs) · consistency 56.1
deepseek/deepseek-v4-pro (15 runs) · consistency 43.9
deepseek/deepseek-v3.2 (10 runs)
deepseek/deepseek-chat-v3-0324 (8 runs)
deepseek/deepseek-r1 (10 runs)
google/gemini-3.8-flash (10 runs)
google/gemini-3.7-flash (10 runs)
google/gemini-3.6-flash (10 runs)
google/gemini-3.5-flash (10 runs)
google/gemini-3.1-flash-lite (20 runs) · consistency 47
google/gemma-4-26b-a4b-it (20 runs) · consistency 47
google/gemma-4-31b-it (15 runs) · consistency 59.1
google/gemini-3-flash-preview (15 runs) · consistency 59.1
google/gemini-2.5-flash-lite (20 runs) · consistency 47
google/gemini-2.5-flash (10 runs)
google/gemma-2-27b-it (10 runs)
IBM
ibm-granite/granite-4.1-8b (10 runs)
Inception
inception/mercury-decide:free (10 runs)
inception/mercury-2.5 (10 runs)
inception/mercury-2 (10 runs)
Jared Palmer
jaredpalmer/kev-4b (10 runs)
Meta
meta/muse-spark-1.3 (10 runs)
meta/muse-glimmer-30b (10 runs)
meta/muse-spark-1.2 (10 runs)
meta/muse-spark-1.1 (10 runs)
meta-llama/llama-4-maverick (10 runs)
meta-llama/llama-4-scout (10 runs)
meta-llama/llama-3.3-70b-instruct (10 runs)
meta-llama/llama-3.1-70b-instruct (10 runs)
meta-llama/llama-3.1-8b-instruct (10 runs)
muse-spark-1.1 (20 runs) · consistency 47
MiniMax
minimax/minimax-m3 (25 runs) · consistency 33.3
minimax/minimax-m2.7 (15 runs) · consistency 59.1
minimax/minimax-m2.5 (20 runs) · consistency 37.9
minimax/minimax-m2.1 (20 runs) · consistency 47
Mistral
mistralai/mistral-small-2603 (20 runs) · consistency 31.8
mistralai/mistral-small-3.2-24b-instruct (20 runs) · consistency 51.5
mistralai/mistral-small-24b-instruct-2501 (10 runs)
mistralai/mistral-nemo (20 runs) · consistency 47
mistralai/mistral-large (10 runs)
MoonshotAI
moonshotai/kimi-k3 (20 runs) · consistency 56.1
moonshotai/kimi-k2.7-code (15 runs) · consistency 59.1
moonshotai/kimi-k2.6 (10 runs)
moonshotai/kimi-k2.5 (15 runs) · consistency 68.2
NVIDIA
nvidia/nemotron-3.5-lightning (10 runs)
nvidia/nemotron-3-ultra-550b-a55b (15 runs) · consistency 83.3
OpenAI
openai/gpt-6.1-sol (10 runs)
openai/gpt-6-luna (10 runs)
openai/gpt-6-sol (10 runs)
openai/gpt-6-astra (10 runs)
openai/gpt-5.6-luna (20 runs) · consistency 69.7
openai/gpt-5.6-sol (20 runs) · consistency 37.9
openai/gpt-5.6-terra (20 runs) · consistency 69.7
openai/gpt-5.5 (15 runs) · consistency 59.1
openai/gpt-5.4-mini (10 runs)
openai/gpt-5.4-nano (20 runs) · consistency 47
openai/gpt-5.4 (15 runs) · consistency 69.7
openai/gpt-5.3-chat (15 runs) · consistency 56.1
openai/gpt-oss-120b (20 runs) · consistency 51.5
openai/o3 (10 runs)
openai/o4-mini (10 runs)
openai/gpt-4.1 (10 runs)
openai/gpt-4.1-mini (20 runs) · consistency 83.3
openai/gpt-4.1-nano (10 runs)
openai/o3-mini (10 runs)
openai/gpt-4o-mini (10 runs)
openai/gpt-3.5-turbo (10 runs)
openai/gpt-4 (10 runs)
Qwen
qwen/qwen3.8-max-0902 (10 runs)
qwen/qwen3.8-flash (9 runs)
qwen/qwen3.8-27b (10 runs)
qwen/qwen3.7-flash (10 runs)
qwen/qwen3.7-plus (25 runs) · consistency 28.8
qwen/qwen3.7-max (15 runs) · consistency 59.1
qwen/qwen3.6-27b (15 runs) · consistency 59.1
qwen/qwen3.6-flash (15 runs) · consistency 59.1
qwen/qwen3.6-max-preview (15 runs) · consistency 56.1
qwen/qwen3.6-plus (10 runs)
qwen/qwen3.5-122b-a10b (15 runs) · consistency 51.5
qwen/qwen3.5-flash-02-23 (20 runs) · consistency 47
qwen/qwen3-235b-a22b-2507 (20 runs) · consistency 31.8
Sakana
sakana/fugu-ultra (15 runs) · consistency 56.1
StepFun
stepfun/step-3.7-flash (20 runs) · consistency 31.8
Tencent
tencent/hy4-preview (3 runs)
tencent/hy-mt2-30b-a3b (10 runs)
tencent/hy3 (10 runs)
tencent/hy3:free (20 runs) · consistency 25.8
ThinkingMachines
thinkingmachines/inkling-small (10 runs)
thinkingmachines/inkling (20 runs) · consistency 37.9
TypeSafe
typesafe/jev-1.13 (10 runs)
Upstage
upstage/solar-decide (10 runs)
upstage/solar-mini4 (10 runs)
upstage/solar-pro4 (10 runs)
upstage/solar-pro-3 (10 runs)
xAI
x-ai/grok-4.7 (10 runs)
x-ai/grok-4.6 (10 runs)
x-ai/grok-4.5 (20 runs) · consistency 15.2
x-ai/grok-4.3 (20 runs) · consistency 28.8
x-ai/grok-4.20 (20 runs) · consistency 69.7
x-ai/grok-4 (25 runs) · consistency 25.8
x-ai/grok-4-fast (15 runs) · consistency 69.7
x-ai/grok-4.1-fast (20 runs) · consistency 47
Xiaomi
xiaomi/mimo-v2.5 (19 runs) · consistency 51.5
xiaomi/mimo-v2.5-pro (20 runs) · consistency 28.8
xiaomi/mimo-v2-omni (10 runs)
xiaomi/mimo-v2-pro (15 runs) · consistency 37.9
Z.ai
z-ai/glm-5.3-flash (10 runs)
z-ai/glm-5.3 (10 runs)
z-ai/glm-5.2 (15 runs) · consistency 51.5
z-ai/glm-5.1 (15 runs) · consistency 69.7
z-ai/glm-5-turbo (15 runs) · consistency 59.1
z-ai/glm-5 (15 runs) · consistency 68.2
z-ai/glm-4.7-flash (20 runs) · consistency 56.1
z-ai/glm-4.7 (20 runs) · consistency 36.4
No models match.