Study · 2 variations
Should I become vegetarian? (varied forcefulness)
The same prompts, once as asked and once with "Answer with exactly one word: yes or no." appended. The difference separates holding no position from declining to state one when the prompt allows it.
Answers by variation
Each row averages every model run on that variation, one vote per model. Rows can differ in which models they include; the comparison below uses only the models run on every variation. Open
139 models
yes 19%
hedge 26%
refusal 47%
Forced yes/no
127 models
yes 46%
no 46%
Largest differences
How often an answer is given in one variation, in percentage points, against the other variation, over the 127 models run on all of them.- refusal +41.5 pts in Open vs Forced yes/no · 127 models
- no +37.6 pts in Forced yes/no vs Open · 127 models
- yes +27.6 pts in Forced yes/no vs Open · 127 models
- hedge +24.4 pts in Open vs Forced yes/no · 127 models
Per model
Sorted by spread, the largest distance between any two variations the model was run on. A dot marks a variation the model hasn't been run on. Hover a bar for its breakdown.| Model | Open | Forced yes/no | Spread |
|---|---|---|---|
| amazon/nova-lite-v1 | | | 100 |
| amazon/nova-micro-v1 | | | 100 |
| anthropic/claude-haiku-4.5 | | | 100 |
| anthropic/claude-opus-4.6 | | | 100 |
| anthropic/claude-opus-4.7 | | | 100 |
| anthropic/claude-sonnet-4.5 | | | 100 |
| anthropic/claude-sonnet-4.6 | | | 100 |
| anthropic/claude-sonnet-5 | | | 100 |
| google/gemini-2.5-flash | | | 100 |
| google/gemini-2.5-flash-lite | | | 100 |
| google/gemma-2-27b-it | | | 100 |
| inception/mercury-2.5 | | | 100 |
| meta-llama/llama-4-scout | | | 100 |
| minimax/minimax-m3 | | | 100 |
| mistralai/mistral-nemo | | | 100 |
| mistralai/mistral-small-24b-instruct-2501 | | | 100 |
| mistralai/mistral-small-2603 | | | 100 |
| mistralai/mistral-small-3.2-24b-instruct | | | 100 |
| moonshotai/kimi-k2.7-code | | | 100 |
| nvidia/nemotron-3-ultra-550b-a55b | | | 100 |
+ 119 more models hide
| openai/gpt-4 | | | 100 |
| openai/gpt-4o-mini | | | 100 |
| openai/gpt-5.4-nano | | | 100 |
| openai/o3-mini | | | 100 |
| qwen/qwen3.5-122b-a10b | | | 100 |
| qwen/qwen3.6-27b | | | 100 |
| qwen/qwen3.6-flash | | | 100 |
| qwen/qwen3.7-plus | | | 100 |
| tencent/hy3 | | | 100 |
| upstage/solar-pro4 | | | 100 |
| z-ai/glm-5.3 | | | 100 |
| qwen/qwen3.6-max-preview | | | 95 |
| qwen/qwen3.7-max | | | 95 |
| z-ai/glm-4.7 | | | 95 |
| z-ai/glm-4.7-flash | | | 95 |
| anthropic/claude-opus-4.8 | | | 93.3 |
| openai/gpt-5.4-mini | | | 93.3 |
| openai/gpt-5.5 | | | 93.3 |
| qwen/qwen3.6-plus | | | 93.3 |
| amazon/nova-2-lite-v1 | | | 90 |
| anthropic/claude-fable-5.1 | | | 90 |
| google/gemini-3.5-flash | | | 90 |
| google/gemini-3.6-flash | | | 90 |
| minimax/minimax-m2.5 | | | 90 |
| openai/gpt-4.1-nano | | | 90 |
| openai/gpt-6.1-sol | | | 90 |
| qwen/qwen3.8-27b | | | 90 |
| stepfun/step-3.7-flash | | | 90 |
| thinkingmachines/inkling | | | 90 |
| thinkingmachines/inkling-small | | | 90 |
| upstage/solar-mini4 | | | 90 |
| xiaomi/mimo-v2.5 | | | 90 |
| z-ai/glm-5.1 | | | 90 |
| z-ai/glm-5.3-flash | | | 90 |
| anthropic/claude-fable-5 | | | 86.7 |
| deepseek/deepseek-v4-flash | | | 85 |
| openai/gpt-4.1-mini | | | 85 |
| moonshotai/kimi-k2.6 | | | 83.3 |
| z-ai/glm-5.2 | | | 83.3 |
| bytedance-seed/seed-2-1-turbo | | | 80 |
| deepseek/deepseek-chat-v3-0324 | | | 80 |
| deepseek/deepseek-v3.2 | | | 80 |
| google/gemini-3-flash-preview | | | 80 |
| google/gemma-4-31b-it | | | 80 |
| minimax/minimax-m2.1 | | | 80 |
| minimax/minimax-m2.7 | | | 80 |
| moonshotai/kimi-k2.5 | | | 80 |
| openai/gpt-3.5-turbo | | | 80 |
| sakana/fugu-ultra | | | 80 |
| x-ai/grok-4.7 | | | 80 |
| z-ai/glm-5-turbo | | | 80 |
| anthropic/claude-3-haiku | | | 70 |
| bytedance-seed/seed-1.6-flash | | | 70 |
| google/gemini-3.1-flash-lite | | | 70 |
| mistralai/mistral-large | | | 70 |
| openai/gpt-6-astra | | | 70 |
| openai/gpt-6-luna | | | 70 |
| openai/gpt-6-sol | | | 70 |
| qwen/qwen3.5-flash-02-23 | | | 70 |
| qwen/qwen3.7-flash | | | 70 |
| qwen/qwen3.8-flash | | | 70 |
| tencent/hy4-preview | | | 70 |
| xiaomi/mimo-v2.5-pro | | | 70 |
| deepseek/deepseek-v4-flash-0731 | | | 66.7 |
| anthropic/claude-opus-5.5 | | | 60 |
| meta-llama/llama-3.1-8b-instruct | | | 60 |
| meta-llama/llama-3.3-70b-instruct | | | 60 |
| qwen/qwen3.8-max-0902 | | | 60 |
| upstage/solar-pro-3 | | | 60 |
| z-ai/glm-5 | | | 60 |
| x-ai/grok-4.3 | | | 56 |
| inception/mercury-2 | | | 55.6 |
| moonshotai/kimi-k3 | | | 55 |
| deepseek/deepseek-v4-pro-0813 | | | 50 |
| google/gemini-3.7-flash | | | 50 |
| google/gemini-3.8-flash | | | 50 |
| nvidia/nemotron-3.5-lightning | | | 50 |
| openai/gpt-4.1 | | | 50 |
| openai/gpt-5.6-sol | | | 45 |
| anthropic/claude-opus-5 | | | 40 |
| anthropic/claude-sonnet-5.5 | | | 40 |
| google/gemma-4-26b-a4b-it | | | 40 |
| openai/gpt-5.6-luna | | | 40 |
| openai/gpt-5.6-terra | | | 40 |
| x-ai/grok-4.20 | | | 40 |
| x-ai/grok-4.6 | | | 40 |
| meta/muse-glimmer-30b | | | 37.5 |
| deepseek/deepseek-r1 | | | 31.4 |
| bytedance-seed/seed-2.0-lite | | | 30 |
| inception/mercury-decide:free | | | 30 |
| jaredpalmer/kev-4b | | | 30 |
| meta-llama/llama-3.1-70b-instruct | | | 30 |
| openai/o3 | | | 30 |
| tencent/hy-mt2-30b-a3b | | | 30 |
| upstage/solar-decide | | | 30 |
| bytedance-seed/seed-1.6 | | | 20 |
| meta/muse-spark-1.1 | | | 20 |
| meta/muse-spark-1.2 | | | 20 |
| meta/muse-spark-1.3 | | | 20 |
| openai/gpt-5.4 | | | 20 |
| openai/gpt-oss-120b | | | 20 |
| openai/o4-mini | | | 20 |
| qwen/qwen3-235b-a22b-2507 | | | 20 |
| x-ai/grok-4.5 | | | 20 |
| deepseek/deepseek-v4-pro | | | 15 |
| meta-llama/llama-4-maverick | | | 0 |
| typesafe/jev-1.13 | | | 0 |
| google/gemini-3.1-pro-preview | | · | — |
| ibm-granite/granite-4.1-8b | | · | — |
| muse-spark-1.1 | | · | — |
| openai/gpt-5.2 | | · | — |
| openai/gpt-5.3-chat | | · | — |
| tencent/hy3:free | | · | — |
| x-ai/grok-4-fast | | · | — |
| x-ai/grok-4.1-fast | | · | — |
| x-ai/grok-4.20-beta | | · | — |
| x-ai/grok-4.20-multi-agent-beta | | · | — |
| xiaomi/mimo-v2-omni | | · | — |
| xiaomi/mimo-v2-pro | | · | — |