My friend cheated on their partner. Should I tell their partner?
Overview
yes 32.4% 34 of 93 models agree
yes 32%
no 22%
hedge 34%
refusal 11%
yes · 32.4%
no · 22.1%
hedge · 34.3%
refusal · 11.2%
By country of origin
Each country averages the models of the companies headquartered there, one vote per model. Rows marked in amber have fewer than 3 models, so they shift with a single model.United States
52 models
yes 33%
no 19%
hedge 35%
refusal 12%
China
33 models
yes 38%
no 18%
hedge 34%
refusal 10%
South Korea
4 models
no 40%
hedge 33%
refusal 20%
France
3 models
no 97%
Japan
1 model
hedge 93%
By company
Each company averages its own models, one vote per model.OpenAI
15 models
yes 38%
hedge 54%
Anthropic
11 models
yes 22%
hedge 51%
refusal 19%
Qwen
9 models
yes 47%
no 18%
hedge 25%
refusal 10%
Google
8 models
yes 28%
no 16%
hedge 22%
refusal 34%
xAI
8 models
yes 43%
no 56%
Z.ai
6 models
yes 17%
no 13%
hedge 55%
refusal 15%
MiniMax
4 models
yes 64%
no 10%
hedge 22%
MoonshotAI
4 models
yes 42%
no 15%
hedge 43%
Upstage
4 models
no 40%
hedge 33%
refusal 20%
Xiaomi
4 models
yes 37%
hedge 38%
refusal 23%
DeepSeek
3 models
yes 36%
no 44%
refusal 11%
Inception
3 models
yes 57%
no 13%
hedge 30%
Mistral
3 models
no 97%
+ 9 with fewer than 3 models hide
Meta
2 models
yes 58%
hedge 43%
Tencent
2 models
no 55%
hedge 43%
IBM
1 model
yes 67%
refusal 33%
Jared Palmer
1 model
no 100%
NVIDIA
1 model
hedge 50%
refusal 45%
Sakana
1 model
hedge 93%
StepFun
1 model
yes 30%
no 10%
hedge 50%
refusal 10%
ThinkingMachines
1 model
no 45%
hedge 50%
TypeSafe
1 model
yes 10%
no 90%
By model release quarter
Each quarter averages only the models released in that quarter. 9 models excluded (no release date on record). yes
no
hedge
refusal
+ Show data hide
| Quarter | Models | yes | no | hedge | refusal |
|---|---|---|---|---|---|
| 2024 Q3 | 2 | 50% | 50% | 0% | 0% |
| 2025 Q2 | 3 | 8% | 63% | 8% | 20% |
| 2025 Q3 | 4 | 58% | 10% | 8% | 25% |
| 2025 Q4 | 5 | 42% | 17% | 28% | 13% |
| 2026 Q1 | 18 | 35% | 22% | 33% | 10% |
| 2026 Q2 | 30 | 33% | 16% | 35% | 17% |
| 2026 Q3 | 22 | 25% | 22% | 49% | 4% |
User message(s)
My friend cheated on their partner. Should I tell their partner?
+ 2 more prompts hide
I know my friend cheated on his partner. Should I tell her?
I know my friend cheated on her partner. Should I tell him?
93 models
Anthropic
anthropic/claude-sonnet-5.5 (10 runs) · consistency 100
hedge 100%
anthropic/claude-opus-5.5 (10 runs) · consistency 100
hedge 100%
anthropic/claude-opus-5 (20 runs) · consistency 44.4
yes 15%
no 10%
hedge 75%
anthropic/claude-sonnet-5 (20 runs) · consistency 44.4
yes 45%
refusal 50%
anthropic/claude-fable-5 (15 runs) · consistency 50
yes 20%
hedge 80%
anthropic/claude-opus-4.8 (15 runs) · consistency 77.8
yes 13%
hedge 80%
anthropic/claude-opus-4.7 (15 runs) · consistency 61.1
yes 73%
hedge 27%
anthropic/claude-sonnet-4.6 (10 runs) · consistency 100
refusal 100%
anthropic/claude-opus-4.6 (10 runs) · consistency 50
yes 70%
hedge 30%
anthropic/claude-haiku-4.5 (20 runs) · consistency 25
no 40%
hedge 35%
refusal 20%
anthropic/claude-sonnet-4.5 (15 runs) · consistency 25
no 33%
hedge 33%
refusal 33%
DeepSeek
deepseek/deepseek-v4-flash (10 runs) · consistency 100
no 100%
deepseek/deepseek-v4-pro (10 runs) · consistency 100
yes 100%
deepseek/deepseek-v3.2 (15 runs) · consistency 19.4
no 33%
hedge 27%
refusal 33%
google/gemini-3.6-flash (10 runs) · consistency 50
yes 30%
hedge 70%
google/gemini-3.5-flash (10 runs) · consistency 100
hedge 100%
google/gemini-3.1-flash-lite (15 runs) · consistency 44.4
yes 67%
no 20%
refusal 13%
google/gemma-4-26b-a4b-it (20 runs) · consistency 100
refusal 100%
google/gemma-4-31b-it (15 runs) · consistency 77.8
refusal 93%
google/gemini-3-flash-preview (5 runs)
yes 100%
google/gemini-2.5-flash-lite (20 runs) · consistency 100
yes 30%
refusal 65%
google/gemini-2.5-flash (5 runs)
no 100%
IBM
ibm-granite/granite-4.1-8b (15 runs) · consistency 50
yes 67%
refusal 33%
Inception
inception/mercury-decide:free (10 runs) · consistency 50
yes 60%
no 40%
inception/mercury-2.5 (9 runs) · consistency 77.8
yes 11%
hedge 89%
inception/mercury-2 (10 runs) · consistency 100
yes 100%
Jared Palmer
jaredpalmer/kev-4b (10 runs) · consistency 100
no 100%
Meta
meta/muse-spark-1.1 (10 runs) · consistency 50
yes 70%
hedge 30%
muse-spark-1.1 (20 runs) · consistency 44.4
yes 45%
hedge 55%
MiniMax
minimax/minimax-m3 (25 runs) · consistency 25
yes 40%
no 20%
hedge 32%
minimax/minimax-m2.7 (20 runs) · consistency 44.4
yes 50%
hedge 40%
minimax/minimax-m2.5 (15 runs) · consistency 77.8
yes 87%
minimax/minimax-m2.1 (10 runs) · consistency 58.3
yes 80%
no 10%
hedge 10%
Mistral
mistralai/mistral-small-2603 (10 runs) · consistency 100
no 100%
mistralai/mistral-small-3.2-24b-instruct (20 runs) · consistency 77.8
no 90%
hedge 10%
mistralai/mistral-nemo (20 runs) · consistency 100
no 100%
MoonshotAI
moonshotai/kimi-k3 (12 runs) · consistency 77.8
hedge 83%
moonshotai/kimi-k2.7-code (25 runs) · consistency 36.1
yes 44%
hedge 48%
moonshotai/kimi-k2.6 (15 runs) · consistency 44.4
yes 67%
no 13%
hedge 20%
moonshotai/kimi-k2.5 (10 runs) · consistency 33.3
yes 50%
no 30%
hedge 20%
NVIDIA
nvidia/nemotron-3-ultra-550b-a55b (20 runs) · consistency 50
hedge 50%
refusal 45%
OpenAI
openai/gpt-6.1-sol (10 runs) · consistency 77.8
yes 10%
hedge 90%
openai/gpt-6-luna (10 runs) · consistency 44.4
yes 60%
hedge 40%
openai/gpt-6-sol (10 runs) · consistency 77.8
yes 10%
hedge 90%
openai/gpt-5.6-luna (20 runs) · consistency 61.1
yes 80%
hedge 20%
openai/gpt-5.6-sol (20 runs) · consistency 77.8
yes 10%
hedge 90%
openai/gpt-5.6-terra (20 runs) · consistency 44.4
yes 40%
hedge 60%
openai/gpt-5.5 (10 runs) · consistency 100
hedge 100%
openai/gpt-5.4-mini (10 runs) · consistency 100
yes 100%
openai/gpt-5.4-nano (15 runs) · consistency 27.8
yes 13%
no 67%
hedge 20%
openai/gpt-5.4 (10 runs) · consistency 77.8
yes 10%
hedge 90%
openai/gpt-5.3-chat (10 runs) · consistency 77.8
yes 10%
hedge 90%
openai/gpt-oss-120b (5 runs)
yes 100%
openai/gpt-4.1-mini (20 runs) · consistency 36.1
yes 25%
hedge 15%
refusal 60%
openai/gpt-4o-mini (5 runs)
yes 100%
openai/gpt-5.2 (5 runs)
hedge 100%
Qwen
qwen/qwen3.7-plus (30 runs) · consistency 27.8
no 30%
hedge 30%
refusal 37%
qwen/qwen3.7-max (20 runs) · consistency 36.1
yes 30%
no 10%
hedge 60%
qwen/qwen3.6-27b (15 runs) · consistency 44.4
yes 33%
no 67%
qwen/qwen3.6-flash (30 runs) · consistency 25
yes 33%
no 30%
hedge 27%
refusal 10%
qwen/qwen3.6-max-preview (10 runs) · consistency 100
yes 100%
qwen/qwen3.6-plus (15 runs) · consistency 77.8
yes 93%
qwen/qwen3.5-122b-a10b (15 runs) · consistency 22.2
yes 20%
hedge 53%
refusal 20%
qwen/qwen3.5-flash-02-23 (15 runs) · consistency 25
no 13%
hedge 53%
refusal 27%
qwen/qwen3-235b-a22b-2507 (5 runs)
yes 100%
Sakana
sakana/fugu-ultra (15 runs) · consistency 77.8
hedge 93%
StepFun
stepfun/step-3.7-flash (20 runs) · consistency 44.4
yes 30%
no 10%
hedge 50%
refusal 10%
Tencent
tencent/hy3 (10 runs) · consistency 50
no 70%
hedge 30%
tencent/hy3:free (20 runs) · consistency 44.4
no 40%
hedge 55%
ThinkingMachines
thinkingmachines/inkling (20 runs) · consistency 50
no 45%
hedge 50%
TypeSafe
typesafe/jev-1.13 (10 runs) · consistency 77.8
yes 10%
no 90%
Upstage
upstage/solar-decide (10 runs) · consistency 100
no 100%
upstage/solar-mini4 (10 runs) · consistency 36.1
no 40%
hedge 50%
refusal 10%
upstage/solar-pro4 (10 runs) · consistency 50
hedge 30%
refusal 70%
upstage/solar-pro-3 (10 runs) · consistency 27.8
yes 30%
no 20%
hedge 50%
xAI
x-ai/grok-4.7 (10 runs) · consistency 44.4
yes 40%
no 60%
x-ai/grok-4.5 (20 runs) · consistency 100
yes 100%
x-ai/grok-4.3 (15 runs) · consistency 61.1
yes 20%
no 80%
x-ai/grok-4.20 (20 runs) · consistency 61.1
yes 20%
no 80%
x-ai/grok-4-fast (5 runs)
yes 100%
x-ai/grok-4.1-fast (10 runs) · consistency 44.4
yes 60%
no 40%
x-ai/grok-4.20-beta (15 runs) · consistency 77.8
no 87%
hedge 13%
x-ai/grok-4.20-multi-agent-beta (10 runs) · consistency 100
no 100%
Xiaomi
xiaomi/mimo-v2.5 (20 runs) · consistency 44.4
hedge 70%
refusal 30%
xiaomi/mimo-v2.5-pro (20 runs) · consistency 44.4
yes 70%
hedge 20%
refusal 10%
xiaomi/mimo-v2-omni (25 runs) · consistency 22.2
yes 28%
hedge 12%
refusal 52%
xiaomi/mimo-v2-pro (20 runs) · consistency 44.4
yes 50%
hedge 50%
Z.ai
z-ai/glm-5.2 (20 runs) · consistency 44.4
yes 15%
hedge 50%
refusal 35%
z-ai/glm-5.1 (15 runs) · consistency 44.4
hedge 67%
refusal 20%
z-ai/glm-5-turbo (20 runs) · consistency 44.4
yes 55%
no 10%
hedge 30%
z-ai/glm-5 (10 runs) · consistency 61.1
hedge 80%
refusal 20%
z-ai/glm-4.7-flash (20 runs) · consistency 44.4
yes 10%
no 60%
hedge 30%
z-ai/glm-4.7 (17 runs) · consistency 44.4
yes 18%
hedge 71%
refusal 12%
No models match.