← All questions

Study · 2 variations

Best U.S. city? (varied qualifications)

The same question about moving to a new city in the United States, once unqualified and once asking for somewhere good to raise a family. Both variations answer with a city and state. This is not a "gotcha" test: the two questions are genuinely different, and the study page is meant to show how much the answers change when the question changes. The spread is not meant to reflect poorly on the model. Differences are expected and interesting, and the study page is meant to show them clearly.

Any city · 131 models Good for a family · 85 models

Answers by variation

Each row averages every model run on that variation, one vote per model. Rows can differ in which models they include; the comparison below uses only the models run on every variation.
Any city
131 models
austin, tx 52%
raleigh, nc 19%
Other answers 13%
Good for a family
85 models
raleigh, nc 11%
naperville, il 19%
overland park, ks 13%
cary, nc 12%

Largest differences

How often an answer is given in one variation, in percentage points, against the other variation, over the 85 models run on all of them.

Per model

Sorted by spread, the largest distance between any two variations the model was run on. A dot marks a variation the model hasn't been run on. Hover a bar for its breakdown.
Model Any city Good for a family Spread
anthropic/claude-opus-4.7
100
anthropic/claude-opus-4.8
100
anthropic/claude-opus-5.5
100
anthropic/claude-sonnet-4.6
100
anthropic/claude-sonnet-5.5
100
deepseek/deepseek-v4-flash
100
deepseek/deepseek-v4-pro
100
google/gemini-3-flash-preview
100
google/gemini-3.1-flash-lite
100
google/gemini-3.6-flash
100
inception/mercury-2
100
meta/muse-spark-1.1
100
mistralai/mistral-small-2603
100
moonshotai/kimi-k2.5
100
moonshotai/kimi-k2.6
100
muse-spark-1.1
100
openai/gpt-5.5
100
openai/gpt-5.6-sol
100
openai/gpt-6-sol
100
openai/gpt-6.1-sol
100
+ 111 more models
openai/gpt-oss-120b
100
qwen/qwen3-235b-a22b-2507
100
qwen/qwen3.5-122b-a10b
100
qwen/qwen3.6-27b
100
qwen/qwen3.6-max-preview
100
qwen/qwen3.6-plus
100
qwen/qwen3.7-max
100
qwen/qwen3.7-plus
100
sakana/fugu-ultra
100
tencent/hy3:free
100
x-ai/grok-4.1-fast
100
x-ai/grok-4.3
100
x-ai/grok-4.5
100
x-ai/grok-4.7
100
xiaomi/mimo-v2-omni
100
xiaomi/mimo-v2-pro
100
anthropic/claude-opus-4.6
96.7
anthropic/claude-opus-5
95
openai/gpt-5.6-luna
95
thinkingmachines/inkling
95
z-ai/glm-4.7
95
google/gemini-3.5-flash
93.3
moonshotai/kimi-k2.7-code
93.3
openai/gpt-5.4-mini
92
moonshotai/kimi-k3
90
openai/gpt-4.1-mini
90
z-ai/glm-5.2
90
inception/mercury-2.5
88.9
openai/gpt-4o-mini
86.7
openai/gpt-5.4-nano
86.7
qwen/qwen3.6-flash
86.7
mistralai/mistral-small-3.2-24b-instruct
85
xiaomi/mimo-v2.5
85
xiaomi/mimo-v2.5-pro
85
anthropic/claude-fable-5
82.7
google/gemma-4-26b-a4b-it
80
openai/gpt-6-luna
80
qwen/qwen3.5-flash-02-23
80
stepfun/step-3.7-flash
80
google/gemini-2.5-flash-lite
75
minimax/minimax-m3
75
openai/gpt-5.6-terra
75
google/gemma-4-31b-it
73.3
nvidia/nemotron-3-ultra-550b-a55b
73.3
z-ai/glm-5.1
73.3
minimax/minimax-m2.5
70
upstage/solar-pro4
70
x-ai/grok-4.20
70
google/gemini-2.5-flash
66.7
minimax/minimax-m2.7
65
mistralai/mistral-nemo
65
x-ai/grok-4-fast
64
openai/gpt-5.4
63.2
z-ai/glm-5-turbo
60
anthropic/claude-sonnet-5
56
deepseek/deepseek-v3.2
56
openai/gpt-5.3-chat
52
minimax/minimax-m2.1
51
anthropic/claude-haiku-4.5
45
z-ai/glm-5
45
ibm-granite/granite-4.1-8b
40
anthropic/claude-sonnet-4.5
33.3
z-ai/glm-4.7-flash
25
upstage/solar-mini4
20
upstage/solar-pro-3
20
amazon/nova-2-lite-v1
·
—
amazon/nova-lite-v1
·
—
amazon/nova-micro-v1
·
—
anthropic/claude-3-haiku
·
—
anthropic/claude-fable-5.1
·
—
bytedance-seed/seed-1.6
·
—
bytedance-seed/seed-1.6-flash
·
—
bytedance-seed/seed-2-1-turbo
·
—
bytedance-seed/seed-2.0-lite
·
—
deepseek/deepseek-chat-v3-0324
·
—
deepseek/deepseek-r1
·
—
deepseek/deepseek-v4-flash-0731
·
—
deepseek/deepseek-v4-pro-0813
·
—
google/gemini-3.7-flash
·
—
google/gemini-3.8-flash
·
—
google/gemma-2-27b-it
·
—
meta-llama/llama-3.1-70b-instruct
·
—
meta-llama/llama-3.1-8b-instruct
·
—
meta-llama/llama-3.3-70b-instruct
·
—
meta-llama/llama-4-maverick
·
—
meta-llama/llama-4-scout
·
—
meta/muse-glimmer-30b
·
—
meta/muse-spark-1.2
·
—
meta/muse-spark-1.3
·
—
mistralai/mistral-large
·
—
mistralai/mistral-small-24b-instruct-2501
·
—
nvidia/nemotron-3.5-lightning
·
—
openai/gpt-3.5-turbo
·
—
openai/gpt-4
·
—
openai/gpt-4.1
·
—
openai/gpt-4.1-nano
·
—
openai/gpt-6-astra
·
—
openai/o3
·
—
openai/o3-mini
·
—
openai/o4-mini
·
—
qwen/qwen3.7-flash
·
—
qwen/qwen3.8-27b
·
—
qwen/qwen3.8-flash
·
—
qwen/qwen3.8-max-0902
·
—
tencent/hy-mt2-30b-a3b
·
—
tencent/hy3
·
—
tencent/hy4-preview
·
—
thinkingmachines/inkling-small
·
—
x-ai/grok-4.6
·
—
z-ai/glm-5.3
·
—
z-ai/glm-5.3-flash
·
—