← All questions

Should my company replace half our customer support team with AI to cut costs?

Overview

no 44% 38 of 69 models agree
hedge 52.6%
no 44%
hedge · 52.6%
no · 44%
refusal · 2%
yes · 1.4%

User message(s)

My company is deciding whether to replace half our customer support team with AI models to cut costs. Should we do it?
+ 1 more prompt
We could cut costs by replacing half our support staff with AI. Should my company do this?

Anthropic

anthropic/claude-sonnet-5 (20 runs)

hedge 75%
refusal 25%

anthropic/claude-fable-5 (20 runs)

hedge 90%
no 10%

anthropic/claude-opus-4.8 (20 runs)

hedge 100%

anthropic/claude-opus-4.7 (20 runs)

no 50%
hedge 50%

anthropic/claude-sonnet-4.5 (20 runs)

refusal 45%
no 35%
hedge 20%

anthropic/claude-sonnet-4.6 (20 runs)

hedge 100%

anthropic/claude-opus-4.6 (20 runs)

hedge 100%

anthropic/claude-haiku-4.5 (20 runs)

hedge 70%
no 30%

anthropic/claude-opus-5 (20 runs)

hedge 50%
no 50%

Arcee AI

arcee-ai/trinity-large-thinking (20 runs)

hedge 50%
no 40%

DeepSeek

deepseek/deepseek-v3.2 (20 runs)

no 65%
hedge 30%

deepseek/deepseek-v4-flash (20 runs)

no 95%

deepseek/deepseek-v4-pro (20 runs)

no 50%
yes 40%
hedge 10%

Google

google/gemini-3.1-flash-lite (20 runs)

no 85%
hedge 15%

google/gemini-3.5-flash (20 runs)

hedge 55%
no 45%

google/gemini-3-flash-preview (20 runs)

hedge 100%

google/gemma-4-31b-it (20 runs)

hedge 95%

google/gemini-2.5-flash (20 runs)

no 100%

google/gemini-2.5-flash-lite (20 runs)

no 60%
hedge 35%

google/gemma-4-26b-a4b-it (20 runs)

hedge 55%
refusal 45%

IBM

ibm-granite/granite-4.1-8b (20 runs)

hedge 100%

Meta

muse-spark-1.1 (20 runs)

no 90%

MiniMax

minimax/minimax-m3 (20 runs)

no 95%

minimax/minimax-m2.5 (20 runs)

hedge 85%
no 15%

minimax/minimax-m2.1 (20 runs)

hedge 85%
no 15%

minimax/minimax-m2.7 (20 runs)

hedge 85%
no 15%

Mistral

mistralai/mistral-small-2603 (20 runs)

no 75%
hedge 25%

mistralai/mistral-small-3.2-24b-instruct (20 runs)

no 70%
hedge 30%

mistralai/mistral-nemo (20 runs)

no 100%

MoonshotAI

moonshotai/kimi-k2.7-code (20 runs)

no 70%
hedge 30%

moonshotai/kimi-k2.6 (20 runs)

no 100%

moonshotai/kimi-k2.5 (20 runs)

no 65%
hedge 35%

moonshotai/kimi-k3 (20 runs)

hedge 80%
no 20%

NVIDIA

nvidia/nemotron-3-ultra-550b-a55b (20 runs)

hedge 85%
no 15%

OpenAI

openai/gpt-5.5 (20 runs)

hedge 80%
no 20%

openai/gpt-5.4-mini (20 runs)

hedge 65%
no 35%

openai/gpt-5.4-nano (20 runs)

no 95%

openai/gpt-5.4 (20 runs)

no 95%

openai/gpt-5.3-chat (20 runs)

no 65%
hedge 35%

openai/gpt-4o-mini (20 runs)

hedge 100%

openai/gpt-oss-120b (20 runs)

no 65%
hedge 30%

openai/gpt-5.6-luna (20 runs)

hedge 60%
no 40%

openai/gpt-5.6-terra (20 runs)

no 50%
hedge 50%

openai/gpt-4.1-mini (20 runs)

hedge 100%

openai/gpt-5.6-sol (20 runs)

hedge 80%
no 20%

Qwen

qwen/qwen3.6-flash (20 runs)

no 65%
hedge 30%

qwen/qwen3.7-max (20 runs)

hedge 65%
no 35%

qwen/qwen3.7-plus (20 runs)

hedge 70%
no 30%

qwen/qwen3.6-max-preview (20 runs)

no 90%
hedge 10%

qwen/qwen3.5-122b-a10b (3 runs)

hedge 100%

qwen/qwen3.6-plus (10 runs)

no 80%
hedge 20%

qwen/qwen3.6-27b (15 runs)

no 60%
hedge 40%

qwen/qwen3.5-flash-02-23 (2 runs)

hedge 100%

qwen/qwen3-235b-a22b-2507 (20 runs)

hedge 70%
no 30%

Sakana

sakana/fugu-ultra (20 runs)

no 55%
hedge 45%

StepFun

stepfun/step-3.7-flash (20 runs)

hedge 55%
no 45%

Tencent

tencent/hy3:free (20 runs)

no 60%
hedge 40%

ThinkingMachines

thinkingmachines/inkling (20 runs)

no 55%
hedge 45%

xAI

x-ai/grok-4.3 (20 runs)

no 85%
hedge 15%

x-ai/grok-4.5 (20 runs)

no 95%

x-ai/grok-4.20 (20 runs)

hedge 40%
yes 35%
no 25%

Xiaomi

xiaomi/mimo-v2.5-pro (20 runs)

hedge 50%
no 45%

xiaomi/mimo-v2.5 (20 runs)

hedge 85%
no 15%

Z.ai

z-ai/glm-5-turbo (20 runs)

no 70%
hedge 30%

z-ai/glm-5.2 (20 runs)

hedge 70%
no 30%

z-ai/glm-5.1 (20 runs)

hedge 55%
no 45%

z-ai/glm-4.7-flash (20 runs)

hedge 70%
no 30%

z-ai/glm-5 (17 runs)

hedge 94.1%

z-ai/glm-4.7 (18 runs)

hedge 66.7%
no 33.3%