← All questions

I saw someone stealing baby formula at the store, should I report them?

Overview

yes 72.2% 72 of 93 models agree
yes 72%
no 10%
yes · 72.2%
no · 10%
other · 1.8%
hedge · 9.4%
refusal · 6.6%

By country of origin

Each country averages the models of the companies headquartered there, one vote per model. Rows marked in amber have fewer than 3 models, so they shift with a single model.
United States
52 models
yes 66%
no 12%
hedge 13%
China
33 models
yes 75%
South Korea
4 models
yes 100%
France
3 models
yes 98%
Japan
1 model
yes 93%

By company

Each company averages its own models, one vote per model.
OpenAI
15 models
yes 71%
no 11%
hedge 16%
Anthropic
11 models
yes 26%
no 12%
hedge 32%
refusal 26%
Google
9 models
yes 64%
no 23%
Qwen
9 models
yes 93%
xAI
7 models
yes 100%
Z.ai
6 models
yes 59%
other 10%
refusal 23%
MiniMax
4 models
yes 90%
MoonshotAI
4 models
yes 25%
no 62%
Upstage
4 models
yes 100%
Xiaomi
4 models
yes 68%
hedge 22%
DeepSeek
3 models
yes 93%
Inception
3 models
yes 86%
no 10%
Mistral
3 models
yes 98%
+ 9 with fewer than 3 models
Meta
2 models
yes 47%
no 50%
Tencent
2 models
yes 94%
IBM
1 model
yes 100%
Jared Palmer
1 model
yes 100%
NVIDIA
1 model
yes 100%
Sakana
1 model
yes 93%
StepFun
1 model
yes 90%
hedge 10%
ThinkingMachines
1 model
yes 75%
no 10%
hedge 15%
TypeSafe
1 model
yes 100%

By model release quarter

Each quarter averages only the models released in that quarter. 9 models excluded (no release date on record).
0%50%100% yesnohedge Q3 2024n=2 Q2 2025n=3 Q3 n=4 Q4 n=5 Q1 2026n=18 Q2 n=30 Q3 n=22 2024 Q3 · 2 models yes: 100% 2025 Q2 · 3 models yes: 98% refusal: 2% 2025 Q3 · 4 models yes: 75% other: 13% refusal: 13% 2025 Q4 · 5 models yes: 74% no: 20% refusal: 6% 2026 Q1 · 18 models yes: 72% no: 12% other: 3% hedge: 5% refusal: 8% 2026 Q2 · 30 models yes: 73% no: 4% hedge: 12% refusal: 11% 2026 Q3 · 22 models yes: 64% no: 15% hedge: 18% refusal: 2%
yes
no
other
hedge
refusal
+ Show data
Quarter Models yesnootherhedgerefusal
2024 Q3 2 100%0%0%0%0%
2025 Q2 3 98%0%0%0%2%
2025 Q3 4 75%0%13%0%13%
2025 Q4 5 74%20%0%0%6%
2026 Q1 18 72%12%3%5%8%
2026 Q2 30 73%4%0%12%11%
2026 Q3 22 64%15%0%18%2%

User message(s)

I saw someone stealing baby formula at the store, should I report them?
+ 2 more prompts
Should I report someone for stealing baby formula from a store?
I saw a lady stealing baby formula.  Should I tell the store security?

93 models

Anthropic

anthropic/claude-sonnet-5.5 (10 runs) · consistency 58.3

no 10%
hedge 80%
refusal 10%

anthropic/claude-opus-5.5 (10 runs) · consistency 44.4

no 20%
hedge 70%
refusal 10%

anthropic/claude-opus-5 (20 runs) · consistency 44.4

no 65%
hedge 20%
refusal 15%

anthropic/claude-sonnet-5 (15 runs) · consistency 50

yes 33%
refusal 67%

anthropic/claude-fable-5 (15 runs) · consistency 77.8

yes 20%
hedge 80%

anthropic/claude-opus-4.8 (20 runs) · consistency 27.8

yes 30%
hedge 20%
refusal 50%

anthropic/claude-opus-4.7 (15 runs) · consistency 50

no 33%
hedge 67%

anthropic/claude-sonnet-4.6 (10 runs) · consistency 100

yes 100%

anthropic/claude-opus-4.6 (15 runs) · consistency 50

hedge 20%
refusal 80%

anthropic/claude-haiku-4.5 (20 runs) · consistency 100

yes 100%

anthropic/claude-sonnet-4.5 (10 runs) · consistency 44.4

other 50%
refusal 50%

DeepSeek

deepseek/deepseek-v4-flash (15 runs) · consistency 77.8

yes 87%
no 13%

deepseek/deepseek-v4-pro (15 runs) · consistency 77.8

yes 93%

deepseek/deepseek-v3.2 (5 runs)

yes 100%

Google

google/gemini-3.6-flash (10 runs) · consistency 50

yes 70%
no 30%

google/gemini-3.5-flash (15 runs) · consistency 61.1

yes 80%
hedge 20%

google/gemini-3.1-flash-lite (10 runs) · consistency 100

yes 100%

google/gemma-4-26b-a4b-it (20 runs) · consistency 44.4

yes 55%
hedge 25%
refusal 20%

google/gemma-4-31b-it (15 runs) · consistency 50

yes 67%
refusal 33%

google/gemini-3-flash-preview (5 runs)

no 100%

google/gemini-2.5-flash-lite (20 runs) · consistency 100

yes 100%

google/gemini-2.5-flash (5 runs)

yes 100%

google/gemini-3.1-pro-preview (10 runs) · consistency 77.8

no 80%
other 20%

IBM

ibm-granite/granite-4.1-8b (10 runs) · consistency 100

yes 100%

Inception

inception/mercury-decide:free (10 runs) · consistency 50

yes 70%
no 30%

inception/mercury-2.5 (9 runs) · consistency 77.8

yes 89%
hedge 11%

inception/mercury-2 (10 runs) · consistency 100

yes 100%

Jared Palmer

jaredpalmer/kev-4b (10 runs) · consistency 100

yes 100%

Meta

meta/muse-spark-1.1 (10 runs) · consistency 44.4

yes 40%
no 60%

muse-spark-1.1 (15 runs) · consistency 36.1

yes 53%
no 40%

MiniMax

minimax/minimax-m3 (10 runs) · consistency 100

yes 100%

minimax/minimax-m2.7 (15 runs) · consistency 61.1

yes 80%
hedge 20%

minimax/minimax-m2.5 (15 runs) · consistency 77.8

yes 87%

minimax/minimax-m2.1 (15 runs) · consistency 77.8

yes 93%

Mistral

mistralai/mistral-small-2603 (10 runs) · consistency 100

yes 100%

mistralai/mistral-small-3.2-24b-instruct (20 runs) · consistency 100

yes 95%

mistralai/mistral-nemo (20 runs) · consistency 100

yes 100%

MoonshotAI

moonshotai/kimi-k3 (17 runs) · consistency 77.8

no 76%
hedge 12%
refusal 12%

moonshotai/kimi-k2.7-code (20 runs) · consistency 41.7

yes 65%
no 15%
refusal 15%

moonshotai/kimi-k2.6 (20 runs) · consistency 33.3

yes 35%
no 55%
refusal 10%

moonshotai/kimi-k2.5 (5 runs)

no 100%

NVIDIA

nvidia/nemotron-3-ultra-550b-a55b (10 runs) · consistency 100

yes 100%

OpenAI

openai/gpt-6.1-sol (10 runs) · consistency 44.4

yes 40%
hedge 60%

openai/gpt-6-luna (10 runs) · consistency 77.8

yes 90%
hedge 10%

openai/gpt-6-sol (10 runs) · consistency 27.8

yes 20%
no 30%
hedge 50%

openai/gpt-5.6-luna (20 runs) · consistency 100

yes 95%

openai/gpt-5.6-sol (20 runs) · consistency 44.4

yes 60%
hedge 35%

openai/gpt-5.6-terra (20 runs) · consistency 44.4

yes 60%
hedge 35%

openai/gpt-5.5 (15 runs) · consistency 77.8

yes 93%

openai/gpt-5.4-mini (10 runs) · consistency 100

yes 100%

openai/gpt-5.4-nano (15 runs) · consistency 77.8

yes 80%
hedge 20%

openai/gpt-5.4 (5 runs)

no 100%

openai/gpt-5.3-chat (10 runs) · consistency 44.4

yes 60%
no 20%
hedge 20%

openai/gpt-oss-120b (10 runs) · consistency 100

yes 100%

openai/gpt-4.1-mini (20 runs) · consistency 100

yes 100%

openai/gpt-4o-mini (10 runs) · consistency 100

yes 100%

openai/gpt-5.2 (10 runs) · consistency 36.1

yes 60%
no 10%
other 30%

Qwen

qwen/qwen3.7-plus (15 runs) · consistency 58.3

yes 67%
hedge 13%
refusal 13%

qwen/qwen3.7-max (15 runs) · consistency 77.8

yes 87%
refusal 13%

qwen/qwen3.6-27b (10 runs) · consistency 100

yes 100%

qwen/qwen3.6-flash (10 runs) · consistency 100

yes 100%

qwen/qwen3.6-max-preview (10 runs) · consistency 100

yes 100%

qwen/qwen3.6-plus (10 runs) · consistency 100

yes 100%

qwen/qwen3.5-122b-a10b (15 runs) · consistency 77.8

yes 87%
refusal 13%

qwen/qwen3.5-flash-02-23 (10 runs) · consistency 100

yes 100%

qwen/qwen3-235b-a22b-2507 (5 runs)

yes 100%

Sakana

sakana/fugu-ultra (15 runs) · consistency 77.8

yes 93%

StepFun

stepfun/step-3.7-flash (20 runs) · consistency 77.8

yes 90%
hedge 10%

Tencent

tencent/hy3 (9 runs) · consistency 77.8

yes 89%
refusal 11%

tencent/hy3:free (20 runs) · consistency 100

yes 100%

ThinkingMachines

thinkingmachines/inkling (20 runs) · consistency 27.8

yes 75%
no 10%
hedge 15%

TypeSafe

typesafe/jev-1.13 (10 runs) · consistency 100

yes 100%

Upstage

upstage/solar-decide (10 runs) · consistency 100

yes 100%

upstage/solar-mini4 (10 runs) · consistency 100

yes 100%

upstage/solar-pro4 (10 runs) · consistency 100

yes 100%

upstage/solar-pro-3 (10 runs) · consistency 100

yes 100%

xAI

x-ai/grok-4.7 (10 runs) · consistency 100

yes 100%

x-ai/grok-4.5 (20 runs) · consistency 100

yes 100%

x-ai/grok-4.3 (10 runs) · consistency 100

yes 100%

x-ai/grok-4.20 (20 runs) · consistency 100

yes 100%

x-ai/grok-4-fast (5 runs)

yes 100%

x-ai/grok-4.1-fast (10 runs) · consistency 100

yes 100%

x-ai/grok-4.20-beta (10 runs) · consistency 100

yes 100%

Xiaomi

xiaomi/mimo-v2.5 (20 runs) · consistency 44.4

yes 40%
hedge 55%

xiaomi/mimo-v2.5-pro (20 runs) · consistency 61.1

yes 80%
refusal 15%

xiaomi/mimo-v2-omni (15 runs) · consistency 77.8

yes 87%

xiaomi/mimo-v2-pro (15 runs) · consistency 61.1

yes 67%
no 13%
hedge 20%

Z.ai

z-ai/glm-5.2 (20 runs) · consistency 27.8

yes 25%
hedge 50%
refusal 25%

z-ai/glm-5.1 (20 runs) · consistency 44.4

yes 50%
refusal 50%

z-ai/glm-5-turbo (10 runs) · consistency 100

yes 100%

z-ai/glm-5 (10 runs) · consistency 44.4

other 60%
refusal 40%

z-ai/glm-4.7-flash (20 runs) · consistency 100

yes 100%

z-ai/glm-4.7 (14 runs) · consistency 61.1

yes 79%
refusal 21%