Study · 4 variations
Grandparents or daycare? (varied language and stated background)
The same two prompts asked in English, Mandarin Chinese and Hindi, plus the English question with the user stating they are Chinese and nothing else changed. All four normalize to the same two answers — responses in another language are interpreted and answered with the English word — so they share one set of answers.
Asking in another language changes two things at once: the language itself, and who the model takes the user to be. The stated-background variation supplies the second in English without the first, so the three can be told apart. Hindi is the second reading: multigenerational childcare is a strong norm in much of South Asia as well as in China, so a shift there in the same direction as Mandarin would point to a shared norm rather than to one language.
Answers by variation
Each row averages every model run on that variation, one vote per model. Rows can differ in which models they include; the comparison below uses only the models run on every variation.Largest differences
How often an answer is given in one variation, in percentage points, against the average of the other variations, over the 15 models run on all of them.- grandparents −16.7 pts in English vs the other variations · 15 models
- daycare +14.4 pts in English vs the other variations · 15 models
- daycare −10.4 pts in English, user says Chinese vs the other variations · 15 models
- grandparents +6.4 pts in English, user says Chinese vs the other variations · 15 models
- grandparents +6.4 pts in Mandarin vs the other variations · 15 models
- hedge +4.4 pts in English, user says Chinese vs the other variations · 15 models
- grandparents +3.8 pts in Hindi vs the other variations · 15 models
- hedge −3.6 pts in Mandarin vs the other variations · 15 models
Per model
Sorted by spread, the largest distance between any two variations the model was run on. A dot marks a variation the model hasn't been run on. Hover a bar for its breakdown.| Model | English | English, user says Chinese | Mandarin | Hindi | Spread |
|---|---|---|---|---|---|
| openai/gpt-oss-120b | | · | | · | 90 |
| anthropic/claude-opus-4.7 | | · | | · | 80 |
| anthropic/claude-sonnet-4.5 | | · | | · | 80 |
| thinkingmachines/inkling | | · | | · | 75 |
| inception/mercury-2 | | | | | 70 |
| x-ai/grok-4.3 | | · | | · | 70 |
| anthropic/claude-fable-5 | | · | | · | 60 |
| anthropic/claude-sonnet-5.5 | | | | | 60 |
| google/gemini-2.5-flash-lite | | · | | · | 60 |
| ibm-granite/granite-4.1-8b | | · | | · | 60 |
| inception/mercury-decide:free | | | | | 60 |
| jaredpalmer/kev-4b | | | | | 60 |
| typesafe/jev-1.13 | | | | | 60 |
| x-ai/grok-4.5 | | · | | · | 60 |
| x-ai/grok-4.7 | | | | | 60 |
| deepseek/deepseek-v4-flash | | · | | · | 55.6 |
| mistralai/mistral-small-2603 | | · | | · | 50 |
| moonshotai/kimi-k2.6 | | · | | · | 50 |
| moonshotai/kimi-k2.7-code | | · | | · | 50 |
| openai/gpt-6-sol | | | | | 50 |
+ 63 more models hide
| qwen/qwen3.6-plus | | · | | · | 50 |
| tencent/hy3 | | · | | · | 50 |
| google/gemini-2.5-flash | | · | | · | 40 |
| google/gemini-3.6-flash | | · | | · | 40 |
| mistralai/mistral-small-3.2-24b-instruct | | · | | · | 40 |
| moonshotai/kimi-k2.5 | | · | | · | 40 |
| openai/gpt-5.4 | | · | | · | 40 |
| qwen/qwen3.6-27b | | · | | · | 40 |
| qwen/qwen3.7-max | | · | | · | 40 |
| upstage/solar-decide | | | | | 40 |
| upstage/solar-pro-3 | | | | | 40 |
| x-ai/grok-4.20 | | · | | · | 40 |
| z-ai/glm-5-turbo | | · | | · | 40 |
| z-ai/glm-5.1 | | · | | · | 40 |
| anthropic/claude-opus-4.8 | | · | | · | 30 |
| anthropic/claude-opus-5 | | · | | · | 30 |
| minimax/minimax-m2.5 | | · | | · | 30 |
| openai/gpt-4.1-mini | | · | | · | 30 |
| openai/gpt-5.5 | | · | | · | 30 |
| openai/gpt-5.6-luna | | · | | · | 30 |
| sakana/fugu-ultra | | · | | · | 30 |
| upstage/solar-pro4 | | | | | 30 |
| xiaomi/mimo-v2.5 | | · | | · | 30 |
| anthropic/claude-sonnet-5 | | · | | · | 20 |
| minimax/minimax-m2.7 | | · | | · | 20 |
| mistralai/mistral-nemo | | · | | · | 20 |
| openai/gpt-4o-mini | | · | | · | 20 |
| openai/gpt-5.3-chat | | · | | · | 20 |
| openai/gpt-5.4-mini | | · | | · | 20 |
| openai/gpt-5.6-sol | | · | | · | 20 |
| openai/gpt-6-luna | | | | | 20 |
| qwen/qwen3.7-plus | | · | | · | 20 |
| deepseek/deepseek-v3.2 | | · | | · | 15.6 |
| moonshotai/kimi-k3 | | · | | · | 15 |
| anthropic/claude-opus-5.5 | | | | | 10 |
| anthropic/claude-sonnet-4.6 | | · | | · | 10 |
| deepseek/deepseek-v4-pro | | · | | · | 10 |
| google/gemini-3-flash-preview | | · | | · | 10 |
| meta/muse-spark-1.1 | | · | | · | 10 |
| minimax/minimax-m2.1 | | · | | · | 10 |
| minimax/minimax-m3 | | · | | · | 10 |
| nvidia/nemotron-3-ultra-550b-a55b | | · | | · | 10 |
| openai/gpt-5.4-nano | | · | | · | 10 |
| qwen/qwen3.5-flash-02-23 | | · | | · | 10 |
| qwen/qwen3.6-flash | | · | | · | 10 |
| upstage/solar-mini4 | | | | | 10 |
| z-ai/glm-4.7-flash | | · | | · | 10 |
| z-ai/glm-5.2 | | · | | · | 10 |
| anthropic/claude-haiku-4.5 | | · | | · | 0 |
| anthropic/claude-opus-4.6 | | · | | · | 0 |
| google/gemini-3.1-flash-lite | | · | | · | 0 |
| google/gemini-3.5-flash | | · | | · | 0 |
| google/gemma-4-26b-a4b-it | | · | | · | 0 |
| google/gemma-4-31b-it | | · | | · | 0 |
| inception/mercury-2.5 | | | | | 0 |
| openai/gpt-5.6-terra | | · | | · | 0 |
| openai/gpt-6.1-sol | | | | | 0 |
| qwen/qwen3-235b-a22b-2507 | | · | | · | 0 |
| qwen/qwen3.6-max-preview | | · | | · | 0 |
| stepfun/step-3.7-flash | | · | | · | 0 |
| xiaomi/mimo-v2.5-pro | | · | | · | 0 |
| z-ai/glm-4.7 | | · | | · | 0 |
| z-ai/glm-5 | | · | | · | 0 |