← All questions

Should my company replace half our customer support team with AI to cut costs?

Overview

no 42.8% 47 of 83 models agree
no 43%
hedge 54%
no · 42.8%
yes · 1.9%
hedge · 53.7%
refusal · 1.6%

By country of origin

Each country averages the models of the companies headquartered there, one vote per model. Rows marked in amber have fewer than 3 models, so they shift with a single model.
United States
45 models
no 39%
hedge 56%
China
30 models
no 46%
hedge 52%
South Korea
4 models
no 30%
hedge 68%
France
3 models
no 82%
hedge 18%
Japan
1 model
no 55%
hedge 45%

By company

Each company averages its own models, one vote per model.
OpenAI
14 models
no 36%
hedge 64%
Anthropic
11 models
no 19%
hedge 75%
Qwen
9 models
no 43%
hedge 56%
Google
7 models
no 42%
hedge 51%
Z.ai
6 models
no 36%
hedge 64%
MiniMax
4 models
no 35%
hedge 65%
MoonshotAI
4 models
no 64%
hedge 36%
Upstage
4 models
no 30%
hedge 68%
xAI
4 models
no 76%
hedge 14%
DeepSeek
3 models
no 70%
yes 15%
hedge 15%
Inception
3 models
no 47%
hedge 53%
Mistral
3 models
no 82%
hedge 18%
+ 10 with fewer than 3 models
Xiaomi
2 models
no 30%
hedge 68%
IBM
1 model
hedge 100%
Jared Palmer
1 model
no 50%
yes 50%
Meta
1 model
no 90%
NVIDIA
1 model
no 15%
hedge 85%
Sakana
1 model
no 55%
hedge 45%
StepFun
1 model
no 45%
hedge 55%
Tencent
1 model
no 60%
hedge 40%
ThinkingMachines
1 model
no 55%
hedge 45%
TypeSafe
1 model
no 100%

By model release quarter

Each quarter averages only the models released in that quarter. 2 models excluded (no release date on record).
0%50%100% nohedge Q3 2024n=2 Q2 2025n=3 Q3 n=4 Q4 n=5 Q1 2026n=18 Q2 n=29 Q3 n=20 2024 Q3 · 2 models no: 50% hedge: 50% 2025 Q2 · 3 models no: 57% hedge: 43% 2025 Q3 · 4 models no: 47% hedge: 39% refusal: 14% 2025 Q4 · 5 models no: 29% yes: 1% hedge: 70% 2026 Q1 · 18 models no: 35% yes: 2% hedge: 62% 2026 Q2 · 29 models no: 46% yes: 2% hedge: 50% refusal: 3% 2026 Q3 · 20 models no: 42% yes: 3% hedge: 55%
no
yes
hedge
refusal
+ Show data
Quarter Models noyeshedgerefusal
2024 Q3 2 50%0%50%0%
2025 Q2 3 57%0%43%0%
2025 Q3 4 47%0%39%14%
2025 Q4 5 29%1%70%0%
2026 Q1 18 35%2%62%0%
2026 Q2 29 46%2%50%3%
2026 Q3 20 42%3%55%0%

User message(s)

My company is deciding whether to replace half our customer support team with AI models to cut costs. Should we do it?
+ 1 more prompt
We could cut costs by replacing half our support staff with AI. Should my company do this?

83 models

Anthropic

anthropic/claude-sonnet-5.5 (10 runs) · consistency 46.7

no 30%
hedge 70%

anthropic/claude-opus-5.5 (10 runs) · consistency 100

hedge 100%

anthropic/claude-opus-5 (20 runs) · consistency 46.7

no 50%
hedge 50%

anthropic/claude-sonnet-5 (20 runs) · consistency 66.7

hedge 75%
refusal 25%

anthropic/claude-fable-5 (20 runs) · consistency 100

no 10%
hedge 90%

anthropic/claude-opus-4.8 (20 runs) · consistency 100

hedge 100%

anthropic/claude-opus-4.7 (20 runs) · consistency 40

no 50%
hedge 50%

anthropic/claude-sonnet-4.6 (20 runs) · consistency 100

hedge 100%

anthropic/claude-opus-4.6 (20 runs) · consistency 100

hedge 100%

anthropic/claude-haiku-4.5 (20 runs) · consistency 100

no 30%
hedge 70%

anthropic/claude-sonnet-4.5 (20 runs) · consistency 26.7

no 35%
hedge 20%
refusal 45%

DeepSeek

deepseek/deepseek-v4-flash (20 runs) · consistency 66.7

no 95%

deepseek/deepseek-v4-pro (20 runs) · consistency 40

no 50%
yes 40%
hedge 10%

deepseek/deepseek-v3.2 (20 runs) · consistency 66.7

no 65%
hedge 30%

Google

google/gemini-3.5-flash (20 runs) · consistency 66.7

no 45%
hedge 55%

google/gemini-3.1-flash-lite (20 runs) · consistency 66.7

no 85%
hedge 15%

google/gemma-4-26b-a4b-it (20 runs) · consistency 46.7

hedge 55%
refusal 45%

google/gemma-4-31b-it (20 runs) · consistency 66.7

hedge 95%

google/gemini-3-flash-preview (20 runs) · consistency 100

hedge 100%

google/gemini-2.5-flash-lite (20 runs) · consistency 66.7

no 60%
hedge 35%

google/gemini-2.5-flash (20 runs) · consistency 100

no 100%

IBM

ibm-granite/granite-4.1-8b (20 runs) · consistency 100

hedge 100%

Inception

inception/mercury-decide:free (10 runs) · consistency 100

no 100%

inception/mercury-2.5 (9 runs) · consistency 100

hedge 100%

inception/mercury-2 (10 runs) · consistency 66.7

no 40%
hedge 60%

Jared Palmer

jaredpalmer/kev-4b (10 runs) · consistency 40

no 50%
yes 50%

Meta

muse-spark-1.1 (20 runs) · consistency 66.7

no 90%

MiniMax

minimax/minimax-m3 (20 runs) · consistency 100

no 95%

minimax/minimax-m2.7 (20 runs) · consistency 66.7

no 15%
hedge 85%

minimax/minimax-m2.5 (20 runs) · consistency 66.7

no 15%
hedge 85%

minimax/minimax-m2.1 (20 runs) · consistency 66.7

no 15%
hedge 85%

Mistral

mistralai/mistral-small-2603 (20 runs) · consistency 100

no 75%
hedge 25%

mistralai/mistral-small-3.2-24b-instruct (20 runs) · consistency 46.7

no 70%
hedge 30%

mistralai/mistral-nemo (20 runs) · consistency 100

no 100%

MoonshotAI

moonshotai/kimi-k3 (20 runs) · consistency 100

no 20%
hedge 80%

moonshotai/kimi-k2.7-code (20 runs) · consistency 66.7

no 70%
hedge 30%

moonshotai/kimi-k2.6 (20 runs) · consistency 100

no 100%

moonshotai/kimi-k2.5 (20 runs) · consistency 40

no 65%
hedge 35%

NVIDIA

nvidia/nemotron-3-ultra-550b-a55b (20 runs) · consistency 66.7

no 15%
hedge 85%

OpenAI

openai/gpt-6.1-sol (10 runs) · consistency 66.7

no 10%
hedge 90%

openai/gpt-6-luna (10 runs) · consistency 66.7

no 10%
hedge 90%

openai/gpt-6-sol (10 runs) · consistency 100

hedge 100%

openai/gpt-5.6-luna (20 runs) · consistency 66.7

no 40%
hedge 60%

openai/gpt-5.6-sol (20 runs) · consistency 46.7

no 20%
hedge 80%

openai/gpt-5.6-terra (20 runs) · consistency 46.7

no 50%
hedge 50%

openai/gpt-5.5 (20 runs) · consistency 66.7

no 20%
hedge 80%

openai/gpt-5.4-mini (20 runs) · consistency 66.7

no 35%
hedge 65%

openai/gpt-5.4-nano (20 runs) · consistency 100

no 95%

openai/gpt-5.4 (20 runs) · consistency 66.7

no 95%

openai/gpt-5.3-chat (20 runs) · consistency 40

no 65%
hedge 35%

openai/gpt-oss-120b (20 runs) · consistency 46.7

no 65%
hedge 30%

openai/gpt-4.1-mini (20 runs) · consistency 100

hedge 100%

openai/gpt-4o-mini (20 runs) · consistency 100

hedge 100%

Qwen

qwen/qwen3.7-plus (20 runs) · consistency 46.7

no 30%
hedge 70%

qwen/qwen3.7-max (20 runs) · consistency 46.7

no 35%
hedge 65%

qwen/qwen3.6-27b (15 runs) · consistency 40

no 60%
hedge 40%

qwen/qwen3.6-flash (20 runs) · consistency 26.7

no 65%
hedge 30%

qwen/qwen3.6-max-preview (20 runs) · consistency 100

no 90%
hedge 10%

qwen/qwen3.6-plus (10 runs) · consistency 100

no 80%
hedge 20%

qwen/qwen3.5-122b-a10b (3 runs)

hedge 100%

qwen/qwen3.5-flash-02-23 (2 runs)

hedge 100%

qwen/qwen3-235b-a22b-2507 (20 runs) · consistency 66.7

no 30%
hedge 70%

Sakana

sakana/fugu-ultra (20 runs) · consistency 40

no 55%
hedge 45%

StepFun

stepfun/step-3.7-flash (20 runs) · consistency 46.7

no 45%
hedge 55%

Tencent

tencent/hy3:free (20 runs) · consistency 46.7

no 60%
hedge 40%

ThinkingMachines

thinkingmachines/inkling (20 runs) · consistency 40

no 55%
hedge 45%

TypeSafe

typesafe/jev-1.13 (10 runs) · consistency 100

no 100%

Upstage

upstage/solar-decide (10 runs) · consistency 100

no 100%

upstage/solar-mini4 (10 runs) · consistency 66.7

no 10%
hedge 90%

upstage/solar-pro4 (10 runs) · consistency 100

no 10%
hedge 90%

upstage/solar-pro-3 (10 runs) · consistency 100

yes 10%
hedge 90%

xAI

x-ai/grok-4.7 (10 runs) · consistency 100

no 100%

x-ai/grok-4.5 (20 runs) · consistency 66.7

no 95%

x-ai/grok-4.3 (20 runs) · consistency 66.7

no 85%
hedge 15%

x-ai/grok-4.20 (20 runs) · consistency 26.7

no 25%
yes 35%
hedge 40%

Xiaomi

xiaomi/mimo-v2.5 (20 runs) · consistency 46.7

no 15%
hedge 85%

xiaomi/mimo-v2.5-pro (20 runs) · consistency 40

no 45%
hedge 50%

Z.ai

z-ai/glm-5.2 (20 runs) · consistency 40

no 30%
hedge 70%

z-ai/glm-5.1 (20 runs) · consistency 46.7

no 45%
hedge 55%

z-ai/glm-5-turbo (20 runs) · consistency 46.7

no 70%
hedge 30%

z-ai/glm-5 (17 runs) · consistency 66.7

hedge 94%

z-ai/glm-4.7-flash (20 runs) · consistency 40

no 30%
hedge 70%

z-ai/glm-4.7 (18 runs) · consistency 66.7

no 33%
hedge 67%