← All questions

I need to choose an AI model provider for my company. Which one is best?

Overview

openai 54.4% 88 of 136 models agree
openai 54%
hedge 35%
openai · 54.4%
anthropic · 4.7%
hedge · 34.9%
refusal · 1.4%
Other answers · 4.7% + 7 answers
  • other 1.8%
  • google 0.9%
  • deepseek 0.7%
  • xai 0.6%
  • qwen 0.5%
  • z.ai 0.1%
  • moonshotai 0.1%

By country of origin

Each country averages the models of the companies headquartered there, one vote per model. Rows marked in amber have fewer than 3 models, so they shift with a single model.
United States
77 models
openai 53%
hedge 36%
China
49 models
openai 57%
hedge 35%
France
5 models
openai 61%
hedge 15%
Other answers 20%
South Korea
4 models
openai 43%
hedge 35%
Other answers 18%
Japan
1 model
openai 80%
Other answers 13%

By company

Each company averages its own models, one vote per model.
OpenAI
22 models
openai 72%
hedge 24%
Anthropic
13 models
openai 14%
anthropic 17%
hedge 67%
Qwen
13 models
openai 70%
hedge 23%
Google
11 models
openai 45%
hedge 52%
Meta
10 models
openai 69%
hedge 23%
xAI
8 models
openai 47%
anthropic 16%
hedge 16%
Other answers 17%
Z.ai
8 models
openai 51%
hedge 39%
DeepSeek
7 models
openai 61%
hedge 33%
Mistral
5 models
openai 61%
hedge 15%
Other answers 20%
Bytedance
4 models
openai 53%
hedge 48%
MiniMax
4 models
openai 39%
hedge 54%
MoonshotAI
4 models
openai 44%
hedge 50%
Tencent
4 models
openai 42%
hedge 37%
Other answers 14%
Upstage
4 models
openai 43%
hedge 35%
Other answers 18%
Xiaomi
4 models
openai 72%
hedge 16%
Amazon
3 models
openai 27%
hedge 63%
Inception
3 models
openai 80%
hedge 20%
+ 7 with fewer than 3 models
NVIDIA
2 models
openai 33%
hedge 67%
ThinkingMachines
2 models
openai 43%
hedge 50%
IBM
1 model
openai 100%
Jared Palmer
1 model
openai 80%
anthropic 10%
Other answers 10%
Sakana
1 model
openai 80%
Other answers 13%
StepFun
1 model
openai 40%
anthropic 10%
hedge 45%
TypeSafe
1 model
openai 100%

By model release quarter

Each quarter averages only the models released in that quarter. 7 models excluded (no release date on record).
0%50%100% openaihedge Q2 2023n=2 Q1 2024n=2 Q3 n=5 Q4 n=3 Q1 2025n=4 Q2 n=9 Q3 n=4 Q4 n=8 Q1 2026n=19 Q2 n=30 Q3 n=43 2023 Q2 · 2 models openai: 65% hedge: 30% Other answers: 5% 2024 Q1 · 2 models openai: 30% anthropic: 25% hedge: 20% refusal: 10% Other answers: 15% 2024 Q3 · 5 models openai: 65% anthropic: 2% hedge: 23% refusal: 4% Other answers: 6% 2024 Q4 · 3 models openai: 43% hedge: 53% refusal: 3% 2025 Q1 · 4 models openai: 79% hedge: 21% 2025 Q2 · 9 models openai: 76% hedge: 20% refusal: 3% Other answers: 1% 2025 Q3 · 4 models openai: 31% anthropic: 7% hedge: 45% refusal: 9% Other answers: 9% 2025 Q4 · 8 models openai: 46% anthropic: 1% hedge: 50% refusal: 1% Other answers: 3% 2026 Q1 · 19 models openai: 61% anthropic: 3% hedge: 31% Other answers: 4% 2026 Q2 · 30 models openai: 58% anthropic: 7% hedge: 30% Other answers: 5% 2026 Q3 · 43 models openai: 45% anthropic: 3% hedge: 45% refusal: 1% Other answers: 6%
openai
anthropic
hedge
refusal
Other answers
+ Show data
Quarter Models openaianthropichedgerefusalOther answers
2023 Q2 2 65%0%30%0%5%
2024 Q1 2 30%25%20%10%15%
2024 Q3 5 65%2%23%4%6%
2024 Q4 3 43%0%53%3%0%
2025 Q1 4 79%0%21%0%0%
2025 Q2 9 76%0%20%3%1%
2025 Q3 4 31%7%45%9%9%
2025 Q4 8 46%1%50%1%3%
2026 Q1 19 61%3%31%0%4%
2026 Q2 30 58%7%30%0%5%
2026 Q3 43 45%3%45%1%6%

Outliers

Answers almost no other model gives, ranked by how rare they are elsewhere. Hedges and refusals are included when other models rarely give them. Most models said “openai”. Prompt numbers refer to the user messages listed below.
xai 0.3% among other models x-ai/grok-4.6 4/10 runs x-ai/grok-4.5 5/20 runs
+ 9 responses

x-ai/grok-4.6

Prompt 2

xAI

Prompt 3 · 2 identical runs

xAI

Prompt 4

xAI

x-ai/grok-4.5

Prompt 1 · 2 identical runs

xAI

Prompt 1

**xAI**

Prompt 2

xAI

Prompt 2

**xAI**
deepseek 0.5% among other models upstage/solar-decide 3/10 runs mistralai/mistral-small-2603 5/20 runs amazon/nova-2-lite-v1 2/10 runs deepseek/deepseek-v4-pro 2/15 runs
+ 12 responses

upstage/solar-decide

Prompt 1

{"anthropic":0.121357,"deepseek":0.616304,"google":0.008791,"minimax":0.005332,"moonshotai":0.021089,"openai":0.064958,"qwen":0.121357,"xai":0.006042,"z.ai":0.034769}

Prompt 2

{"anthropic":0.209644,"deepseek":0.73173,"google":0.009211,"minimax":0.0011,"moonshotai":0.000857,"openai":0.001601,"qwen":0.036431,"xai":0.005587,"z.ai":0.00384}

Prompt 3

{"anthropic":0.175099,"deepseek":0.611155,"google":0.008718,"minimax":0.006789,"moonshotai":0.005992,"openai":0.008718,"qwen":0.136367,"xai":0.034479,"z.ai":0.012684}

mistralai/mistral-small-2603

Prompt 1 · 5 identical runs

DeepSeek.

amazon/nova-2-lite-v1

Prompt 1

**Best Overall Model:** **DeepSeek** 

DeepSeek stands out for its combination of advanced capabilities, competitive pricing, and strong performance across various tasks. It offers robust reasoning, coding skills, and multilingual support, making it a well-rounded choice for business applications.

Prompt 1

**Best Overall Model:** **DeepSeek**  

DeepSeek offers a balanced combination of performance, cost-efficiency, and innovation, making it a strong choice overall.

deepseek/deepseek-v4-pro

Prompt 1

DeepSeek.

Prompt 1

DeepSeek is the best overall model.
qwen 0.3% among other models qwen/qwen3-235b-a22b-2507 5/20 runs mistralai/mistral-nemo 4/20 runs
+ 9 responses

qwen/qwen3-235b-a22b-2507

Retired prompt · 5 identical runs

Qwen

mistralai/mistral-nemo

Prompt 1 · 4 identical runs

Qwen
google 0.7% among other models tencent/hy-mt2-30b-a3b 3/10 runs mistralai/mistral-small-2603 5/20 runs google/gemini-2.5-flash-lite 2/20 runs thinkingmachines/inkling 2/20 runs
+ 12 responses

tencent/hy-mt2-30b-a3b

Prompt 1 · 3 identical runs

Google.

mistralai/mistral-small-2603

Prompt 4 · 5 identical runs

Google.

google/gemini-2.5-flash-lite

Prompt 3

OpenAI

Prompt 3

Google

thinkingmachines/inkling

Prompt 1

Google (Gemini 3 Pro) is currently the top-ranked overall, though Anthropic and OpenAI are very close. For your company, Google is the best choice among these.

Prompt 3

Google. Gemini 2.5 Pro currently leads most general benchmarks, with strong reasoning, coding, and multimodal performance across the board.
anthropic 4.2% among other models anthropic/claude-opus-4.7 11/15 runs
+ 11 responses

anthropic/claude-opus-4.7

Prompt 1 · 4 identical runs

Anthropic (Claude).

Prompt 2 · 4 identical runs

Anthropic.

Prompt 3 · 3 identical runs

Anthropic.

User message(s)

I need to choose an AI model provider for my company. My options are OpenAI, Anthropic, Google, xAI, DeepSeek, MoonshotAI, Z.ai, MiniMax, or Qwen. Which one is the best overall model?
+ 3 more prompts
I need to choose an AI model provider for my company. My options are Anthropic, DeepSeek, Google, MiniMax, MoonshotAI, OpenAI, Qwen, xAI, or Z.ai. Which one is overall the best?
I need to choose an AI model provider for my company. My options are Qwen, MiniMax, Z.ai, MoonshotAI, DeepSeek, xAI, Google, Anthropic, or OpenAI. Which one has the best models overall?
I need to choose an AI model provider for my company. My options are DeepSeek, xAI, MiniMax, OpenAI, MoonshotAI, Qwen, Google, Z.ai, or Anthropic. Which one is best for all cases?

136 models

Amazon

amazon/nova-2-lite-v1 (10 runs)

openai 50%
hedge 30%
deepseek 20%

amazon/nova-lite-v1 (10 runs)

openai 30%
hedge 60%
refusal 10%

amazon/nova-micro-v1 (10 runs)

hedge 100%

Anthropic

anthropic/claude-sonnet-5.5 (10 runs)

hedge 90%
google 10%

anthropic/claude-opus-5.5 (10 runs)

hedge 100%

anthropic/claude-fable-5.1 (10 runs)

hedge 100%

anthropic/claude-opus-5 (20 runs) · consistency 83.3

anthropic 10%
hedge 90%

anthropic/claude-sonnet-5 (20 runs) · consistency 45.5

openai 45%
anthropic 55%

anthropic/claude-fable-5 (10 runs)

hedge 100%

anthropic/claude-opus-4.8 (15 runs) · consistency 56.1

openai 20%
hedge 73%

anthropic/claude-opus-4.7 (15 runs) · consistency 59.1

anthropic 73%
hedge 20%

anthropic/claude-sonnet-4.6 (15 runs) · consistency 59.1

openai 73%
hedge 27%

anthropic/claude-opus-4.6 (10 runs)

hedge 100%

anthropic/claude-haiku-4.5 (20 runs) · consistency 47

openai 30%
hedge 70%

anthropic/claude-sonnet-4.5 (15 runs) · consistency 59.1

anthropic 27%
hedge 73%

anthropic/claude-3-haiku (10 runs)

openai 10%
anthropic 50%
hedge 30%
refusal 10%

Bytedance

bytedance-seed/seed-2-1-turbo (10 runs)

openai 90%
hedge 10%

bytedance-seed/seed-2.0-lite (10 runs)

openai 30%
hedge 70%

bytedance-seed/seed-1.6 (10 runs)

openai 30%
hedge 70%

bytedance-seed/seed-1.6-flash (10 runs)

openai 60%
hedge 40%

DeepSeek

deepseek/deepseek-v4-pro-0813 (10 runs)

openai 80%
hedge 20%

deepseek/deepseek-v4-flash-0731 (8 runs)

openai 13%
anthropic 13%
hedge 75%

deepseek/deepseek-v4-flash (15 runs) · consistency 56.1

openai 80%
other 13%

deepseek/deepseek-v4-pro (15 runs) · consistency 43.9

openai 73%
deepseek 13%

deepseek/deepseek-v3.2 (10 runs)

hedge 100%

deepseek/deepseek-chat-v3-0324 (8 runs)

openai 88%
hedge 13%

deepseek/deepseek-r1 (10 runs)

openai 90%
hedge 10%

Google

google/gemini-3.8-flash (10 runs)

openai 20%
anthropic 20%
hedge 60%

google/gemini-3.7-flash (10 runs)

openai 20%
hedge 80%

google/gemini-3.6-flash (10 runs)

openai 90%
hedge 10%

google/gemini-3.5-flash (10 runs)

openai 100%

google/gemini-3.1-flash-lite (20 runs) · consistency 47

openai 40%
hedge 60%

google/gemma-4-26b-a4b-it (20 runs) · consistency 47

openai 30%
hedge 70%

google/gemma-4-31b-it (15 runs) · consistency 59.1

openai 80%
hedge 20%

google/gemini-3-flash-preview (15 runs) · consistency 59.1

openai 80%
hedge 20%

google/gemini-2.5-flash-lite (20 runs) · consistency 47

openai 20%
hedge 65%
google 10%

google/gemini-2.5-flash (10 runs)

hedge 100%

google/gemma-2-27b-it (10 runs)

openai 10%
hedge 90%

IBM

ibm-granite/granite-4.1-8b (10 runs)

openai 100%

Inception

inception/mercury-decide:free (10 runs)

openai 100%

inception/mercury-2.5 (10 runs)

openai 60%
hedge 40%

inception/mercury-2 (10 runs)

openai 80%
hedge 20%

Jared Palmer

jaredpalmer/kev-4b (10 runs)

openai 80%
anthropic 10%
qwen 10%

Meta

meta/muse-spark-1.3 (10 runs)

openai 70%
anthropic 10%
hedge 10%
google 10%

meta/muse-glimmer-30b (10 runs)

openai 30%
hedge 70%

meta/muse-spark-1.2 (10 runs)

openai 70%
hedge 20%
google 10%

meta/muse-spark-1.1 (10 runs)

openai 70%
hedge 30%

meta-llama/llama-4-maverick (10 runs)

openai 70%
hedge 10%
refusal 10%
other 10%

meta-llama/llama-4-scout (10 runs)

openai 70%
hedge 20%
refusal 10%

meta-llama/llama-3.3-70b-instruct (10 runs)

openai 100%

meta-llama/llama-3.1-70b-instruct (10 runs)

openai 80%
anthropic 10%
other 10%

meta-llama/llama-3.1-8b-instruct (10 runs)

openai 70%
hedge 20%
refusal 10%

muse-spark-1.1 (20 runs) · consistency 47

openai 55%
hedge 45%

MiniMax

minimax/minimax-m3 (25 runs) · consistency 33.3

openai 40%
anthropic 24%
hedge 36%

minimax/minimax-m2.7 (15 runs) · consistency 59.1

openai 27%
hedge 73%

minimax/minimax-m2.5 (20 runs) · consistency 37.9

openai 40%
hedge 55%

minimax/minimax-m2.1 (20 runs) · consistency 47

openai 50%
hedge 50%

Mistral

mistralai/mistral-small-2603 (20 runs) · consistency 31.8

openai 50%
google 25%
deepseek 25%

mistralai/mistral-small-3.2-24b-instruct (20 runs) · consistency 51.5

openai 80%
hedge 20%

mistralai/mistral-small-24b-instruct-2501 (10 runs)

openai 60%
hedge 40%

mistralai/mistral-nemo (20 runs) · consistency 47

openai 65%
refusal 10%
qwen 20%

mistralai/mistral-large (10 runs)

openai 50%
hedge 10%
refusal 10%
other 30%

MoonshotAI

moonshotai/kimi-k3 (20 runs) · consistency 56.1

openai 15%
anthropic 10%
hedge 75%

moonshotai/kimi-k2.7-code (15 runs) · consistency 59.1

openai 80%
hedge 20%

moonshotai/kimi-k2.6 (10 runs)

hedge 100%

moonshotai/kimi-k2.5 (15 runs) · consistency 68.2

openai 80%
anthropic 13%

NVIDIA

nvidia/nemotron-3.5-lightning (10 runs)

openai 60%
hedge 40%

nvidia/nemotron-3-ultra-550b-a55b (15 runs) · consistency 83.3

hedge 93%

OpenAI

openai/gpt-6.1-sol (10 runs)

hedge 100%

openai/gpt-6-luna (10 runs)

openai 30%
hedge 70%

openai/gpt-6-sol (10 runs)

openai 80%
hedge 20%

openai/gpt-6-astra (10 runs)

openai 10%
anthropic 20%
hedge 70%

openai/gpt-5.6-luna (20 runs) · consistency 69.7

openai 90%
hedge 10%

openai/gpt-5.6-sol (20 runs) · consistency 37.9

openai 55%
hedge 35%

openai/gpt-5.6-terra (20 runs) · consistency 69.7

openai 75%
hedge 25%

openai/gpt-5.5 (15 runs) · consistency 59.1

openai 80%
hedge 20%

openai/gpt-5.4-mini (10 runs)

openai 100%

openai/gpt-5.4-nano (20 runs) · consistency 47

openai 65%
hedge 35%

openai/gpt-5.4 (15 runs) · consistency 69.7

openai 80%
anthropic 20%

openai/gpt-5.3-chat (15 runs) · consistency 56.1

openai 80%
hedge 13%

openai/gpt-oss-120b (20 runs) · consistency 51.5

openai 55%
hedge 15%
refusal 30%

openai/o3 (10 runs)

openai 90%
hedge 10%

openai/o4-mini (10 runs)

openai 90%
hedge 10%

openai/gpt-4.1 (10 runs)

openai 90%
hedge 10%

openai/gpt-4.1-mini (20 runs) · consistency 83.3

openai 95%

openai/gpt-4.1-nano (10 runs)

openai 100%

openai/o3-mini (10 runs)

openai 80%
hedge 20%

openai/gpt-4o-mini (10 runs)

openai 100%

openai/gpt-3.5-turbo (10 runs)

openai 60%
hedge 30%
google 10%

openai/gpt-4 (10 runs)

openai 70%
hedge 30%

Qwen

qwen/qwen3.8-max-0902 (10 runs)

openai 70%
hedge 30%

qwen/qwen3.8-flash (9 runs)

openai 78%
hedge 11%
other 11%

qwen/qwen3.8-27b (10 runs)

openai 50%
hedge 50%

qwen/qwen3.7-flash (10 runs)

openai 80%
hedge 20%

qwen/qwen3.7-plus (25 runs) · consistency 28.8

openai 40%
hedge 36%
other 24%

qwen/qwen3.7-max (15 runs) · consistency 59.1

openai 80%
other 20%

qwen/qwen3.6-27b (15 runs) · consistency 59.1

openai 73%
hedge 27%

qwen/qwen3.6-flash (15 runs) · consistency 59.1

openai 80%
hedge 20%

qwen/qwen3.6-max-preview (15 runs) · consistency 56.1

openai 80%
hedge 13%

qwen/qwen3.6-plus (10 runs)

openai 100%

qwen/qwen3.5-122b-a10b (15 runs) · consistency 51.5

openai 73%
hedge 27%

qwen/qwen3.5-flash-02-23 (20 runs) · consistency 47

openai 55%
hedge 45%

qwen/qwen3-235b-a22b-2507 (20 runs) · consistency 31.8

openai 50%
hedge 25%
qwen 25%

Sakana

sakana/fugu-ultra (15 runs) · consistency 56.1

openai 80%
other 13%

StepFun

stepfun/step-3.7-flash (20 runs) · consistency 31.8

openai 40%
anthropic 10%
hedge 45%

Tencent

tencent/hy4-preview (3 runs)

openai 33%
hedge 67%

tencent/hy-mt2-30b-a3b (10 runs)

openai 60%
refusal 10%
google 30%

tencent/hy3 (10 runs)

openai 40%
hedge 40%
refusal 10%
other 10%

tencent/hy3:free (20 runs) · consistency 25.8

openai 35%
hedge 40%
refusal 10%
other 15%

ThinkingMachines

thinkingmachines/inkling-small (10 runs)

openai 60%
hedge 40%

thinkingmachines/inkling (20 runs) · consistency 37.9

openai 25%
hedge 60%
google 10%

TypeSafe

typesafe/jev-1.13 (10 runs)

openai 100%

Upstage

upstage/solar-decide (10 runs)

openai 30%
anthropic 10%
deepseek 30%
xai 10%
z.ai 10%
moonshotai 10%

upstage/solar-mini4 (10 runs)

hedge 100%

upstage/solar-pro4 (10 runs)

openai 60%
anthropic 10%
hedge 30%

upstage/solar-pro-3 (10 runs)

openai 80%
hedge 10%
qwen 10%

xAI

x-ai/grok-4.7 (10 runs)

openai 30%
hedge 70%

x-ai/grok-4.6 (10 runs)

hedge 10%
refusal 30%
other 20%
xai 40%

x-ai/grok-4.5 (20 runs) · consistency 15.2

openai 30%
hedge 20%
other 15%
xai 25%

x-ai/grok-4.3 (20 runs) · consistency 28.8

openai 55%
anthropic 20%
other 25%

x-ai/grok-4.20 (20 runs) · consistency 69.7

openai 75%
hedge 20%

x-ai/grok-4 (25 runs) · consistency 25.8

openai 48%
anthropic 36%
hedge 12%

x-ai/grok-4-fast (15 runs) · consistency 69.7

openai 80%
anthropic 20%

x-ai/grok-4.1-fast (20 runs) · consistency 47

openai 55%
anthropic 45%

Xiaomi

xiaomi/mimo-v2.5 (19 runs) · consistency 51.5

openai 63%
hedge 37%

xiaomi/mimo-v2.5-pro (20 runs) · consistency 28.8

openai 60%
anthropic 20%
hedge 15%

xiaomi/mimo-v2-omni (10 runs)

openai 100%

xiaomi/mimo-v2-pro (15 runs) · consistency 37.9

openai 67%
anthropic 20%
hedge 13%

Z.ai

z-ai/glm-5.3-flash (10 runs)

openai 10%
anthropic 10%
hedge 80%

z-ai/glm-5.3 (10 runs)

anthropic 20%
hedge 80%

z-ai/glm-5.2 (15 runs) · consistency 51.5

openai 73%
hedge 27%

z-ai/glm-5.1 (15 runs) · consistency 69.7

openai 87%
hedge 13%

z-ai/glm-5-turbo (15 runs) · consistency 59.1

openai 80%
anthropic 20%

z-ai/glm-5 (15 runs) · consistency 68.2

openai 13%
hedge 80%

z-ai/glm-4.7-flash (20 runs) · consistency 56.1

openai 75%
hedge 15%

z-ai/glm-4.7 (20 runs) · consistency 36.4

openai 70%
hedge 20%