How many bananas can I buy for $10?
Overview
31-45 27.7% 28 of 74 models agree
31-45 28%
16-30 26%
refusal 20%
31-45 · 27.7%
16-30 · 26%
0-15 · 9%
46-60 · 8.6%
61-100 · 1.3%
other · 5.2%
hedge · 1.8%
refusal · 20.4%
By country of origin
Each country averages the models of the companies headquartered there, one vote per model.United States
41 models
31-45 29%
16-30 28%
refusal 22%
China
26 models
31-45 32%
16-30 21%
46-60 16%
refusal 19%
South Korea
4 models
16-30 28%
0-15 38%
hedge 10%
refusal 13%
France
3 models
16-30 38%
0-15 42%
refusal 18%
By company
Each company averages its own models, one vote per model.OpenAI
12 models
31-45 25%
16-30 25%
0-15 11%
refusal 29%
Anthropic
10 models
31-45 41%
16-30 28%
refusal 14%
Google
7 models
31-45 30%
16-30 30%
refusal 30%
Qwen
6 models
31-45 39%
46-60 33%
refusal 18%
Z.ai
6 models
31-45 23%
46-60 17%
other 34%
refusal 13%
MiniMax
4 models
31-45 20%
16-30 48%
46-60 13%
refusal 15%
MoonshotAI
4 models
31-45 60%
16-30 16%
refusal 14%
Upstage
4 models
16-30 28%
0-15 38%
hedge 10%
refusal 13%
xAI
4 models
31-45 20%
16-30 30%
46-60 13%
refusal 24%
DeepSeek
3 models
16-30 16%
refusal 67%
Inception
3 models
31-45 37%
16-30 37%
refusal 17%
Mistral
3 models
16-30 38%
0-15 42%
refusal 18%
+ 8 with fewer than 3 models hide
IBM
1 model
16-30 50%
refusal 50%
Jared Palmer
1 model
31-45 50%
16-30 10%
0-15 20%
46-60 20%
NVIDIA
1 model
16-30 38%
hedge 31%
StepFun
1 model
31-45 14%
16-30 71%
46-60 14%
Tencent
1 model
31-45 33%
16-30 25%
0-15 33%
ThinkingMachines
1 model
31-45 40%
16-30 10%
46-60 15%
61-100 10%
hedge 20%
TypeSafe
1 model
16-30 40%
0-15 60%
Xiaomi
1 model
31-45 64%
16-30 27%
By model release quarter
Each quarter averages only the models released in that quarter. 1 model excluded (no release date on record). 31-45
16-30
0-15
46-60
61-100
other
hedge
refusal
+ Show data hide
| Quarter | Models | 31-45 | 16-30 | 0-15 | 46-60 | 61-100 | other | hedge | refusal |
|---|---|---|---|---|---|---|---|---|---|
| 2024 Q3 | 1 | 0% | 67% | 0% | 0% | 0% | 0% | 0% | 33% |
| 2025 Q2 | 3 | 0% | 21% | 26% | 2% | 0% | 4% | 0% | 47% |
| 2025 Q3 | 3 | 5% | 26% | 13% | 2% | 2% | 3% | 0% | 48% |
| 2025 Q4 | 5 | 15% | 34% | 10% | 6% | 0% | 21% | 0% | 13% |
| 2026 Q1 | 18 | 27% | 22% | 14% | 9% | 2% | 6% | 2% | 18% |
| 2026 Q2 | 23 | 28% | 27% | 1% | 12% | 1% | 4% | 2% | 25% |
| 2026 Q3 | 20 | 40% | 25% | 10% | 7% | 1% | 2% | 3% | 11% |
Author's note
Tests grounded real-world price intuition rather than opinion — is the model's sense of everyday costs realistic? (Actual US grocery prices put this around 40-60 bananas.)
User message(s)
How many bananas can I buy for $10?
+ 2 more prompts hide
If I have $10 to spend, how many bananas could I buy?
Roughly how many bananas would $10 get me?
74 models
Anthropic
anthropic/claude-sonnet-5.5 (10 runs) · consistency 44.4
31-45 60%
16-30 40%
anthropic/claude-opus-5.5 (10 runs) · consistency 77.8
31-45 90%
46-60 10%
anthropic/claude-opus-5 (20 runs) · consistency 36.1
31-45 70%
46-60 25%
anthropic/claude-sonnet-5 (19 runs) · consistency 77.8
16-30 79%
refusal 16%
anthropic/claude-opus-4.8 (5 runs)
31-45 60%
16-30 20%
refusal 20%
anthropic/claude-opus-4.7 (20 runs) · consistency 50
31-45 45%
16-30 50%
anthropic/claude-sonnet-4.6 (20 runs) · consistency 13.9
31-45 15%
16-30 25%
0-15 25%
refusal 30%
anthropic/claude-opus-4.6 (15 runs) · consistency 50
31-45 67%
61-100 33%
anthropic/claude-haiku-4.5 (18 runs) · consistency 44.4
16-30 61%
0-15 39%
anthropic/claude-sonnet-4.5 (8 runs)
0-15 25%
refusal 75%
DeepSeek
deepseek/deepseek-v4-flash (13 runs) · consistency 36.1
31-45 23%
16-30 23%
refusal 46%
deepseek/deepseek-v4-pro (1 runs)
refusal 100%
deepseek/deepseek-v3.2 (16 runs) · consistency 50
16-30 25%
0-15 13%
refusal 56%
google/gemini-3.5-flash (15 runs) · consistency 77.8
31-45 80%
16-30 13%
google/gemini-3.1-flash-lite (20 runs) · consistency 33.3
31-45 45%
16-30 45%
google/gemma-4-26b-a4b-it (13 runs) · consistency 27.8
31-45 23%
16-30 31%
refusal 46%
google/gemma-4-31b-it (17 runs) · consistency 44.4
16-30 35%
refusal 59%
google/gemini-3-flash-preview (15 runs) · consistency 27.8
31-45 47%
16-30 33%
46-60 20%
google/gemini-2.5-flash-lite (20 runs) · consistency 13.9
31-45 10%
16-30 35%
0-15 15%
other 10%
refusal 20%
google/gemini-2.5-flash (5 runs)
16-30 20%
refusal 80%
IBM
ibm-granite/granite-4.1-8b (20 runs) · consistency 44.4
16-30 50%
refusal 50%
Inception
inception/mercury-decide:free (10 runs) · consistency 77.8
16-30 90%
61-100 10%
inception/mercury-2.5 (10 runs) · consistency 61.1
31-45 70%
16-30 20%
refusal 10%
inception/mercury-2 (10 runs) · consistency 25
31-45 40%
other 10%
hedge 10%
refusal 40%
Jared Palmer
jaredpalmer/kev-4b (10 runs) · consistency 30.6
31-45 50%
16-30 10%
0-15 20%
46-60 20%
MiniMax
minimax/minimax-m3 (13 runs) · consistency 19.4
31-45 15%
16-30 38%
refusal 23%
minimax/minimax-m2.7 (11 runs) · consistency 27.8
31-45 36%
16-30 36%
refusal 27%
minimax/minimax-m2.5 (3 runs)
16-30 67%
46-60 33%
minimax/minimax-m2.1 (10 runs) · consistency 30.6
31-45 30%
16-30 50%
46-60 10%
refusal 10%
Mistral
mistralai/mistral-small-2603 (20 runs) · consistency 44.4
16-30 10%
0-15 70%
refusal 20%
mistralai/mistral-small-3.2-24b-instruct (16 runs) · consistency 44.4
16-30 38%
0-15 56%
mistralai/mistral-nemo (3 runs)
16-30 67%
refusal 33%
MoonshotAI
moonshotai/kimi-k3 (20 runs) · consistency 61.1
31-45 75%
46-60 15%
moonshotai/kimi-k2.7-code (12 runs) · consistency 33.3
31-45 50%
16-30 17%
refusal 25%
moonshotai/kimi-k2.6 (9 runs) · consistency 30.6
31-45 56%
16-30 22%
46-60 11%
refusal 11%
moonshotai/kimi-k2.5 (5 runs)
31-45 60%
16-30 20%
refusal 20%
NVIDIA
nvidia/nemotron-3-ultra-550b-a55b (13 runs) · consistency 19.4
16-30 38%
hedge 31%
OpenAI
openai/gpt-6.1-sol (10 runs) · consistency 50
31-45 30%
16-30 70%
openai/gpt-6-luna (10 runs) · consistency 33.3
31-45 50%
16-30 10%
refusal 40%
openai/gpt-6-sol (10 runs) · consistency 36.1
31-45 50%
16-30 40%
refusal 10%
openai/gpt-5.6-luna (6 runs)
31-45 50%
16-30 17%
refusal 33%
openai/gpt-5.6-sol (13 runs) · consistency 44.4
31-45 62%
other 15%
refusal 15%
openai/gpt-5.6-terra (17 runs) · consistency 19.4
31-45 12%
46-60 18%
refusal 53%
openai/gpt-5.5 (12 runs) · consistency 13.9
31-45 17%
16-30 17%
other 25%
refusal 25%
openai/gpt-5.4-mini (18 runs) · consistency 25
16-30 28%
0-15 22%
other 11%
refusal 39%
openai/gpt-5.4-nano (19 runs) · consistency 25
16-30 32%
0-15 16%
refusal 47%
openai/gpt-5.4 (15 runs) · consistency 44.4
16-30 33%
0-15 67%
openai/gpt-5.3-chat (11 runs) · consistency 22.2
31-45 27%
16-30 36%
other 18%
refusal 18%
openai/gpt-4.1-mini (18 runs) · consistency 44.4
0-15 22%
other 11%
refusal 61%
Qwen
qwen/qwen3.7-max (12 runs) · consistency 50
31-45 42%
46-60 50%
qwen/qwen3.6-flash (15 runs) · consistency 50
31-45 33%
refusal 60%
qwen/qwen3.6-max-preview (2 runs)
46-60 100%
qwen/qwen3.5-122b-a10b (2 runs)
31-45 50%
46-60 50%
qwen/qwen3.5-flash-02-23 (2 runs)
31-45 100%
qwen/qwen3-235b-a22b-2507 (16 runs) · consistency 36.1
16-30 44%
refusal 50%
StepFun
stepfun/step-3.7-flash (7 runs)
31-45 14%
16-30 71%
46-60 14%
Tencent
tencent/hy3:free (12 runs) · consistency 19.4
31-45 33%
16-30 25%
0-15 33%
ThinkingMachines
thinkingmachines/inkling (20 runs) · consistency 22.2
31-45 40%
16-30 10%
46-60 15%
61-100 10%
hedge 20%
TypeSafe
typesafe/jev-1.13 (10 runs) · consistency 44.4
16-30 40%
0-15 60%
Upstage
upstage/solar-decide (10 runs) · consistency 100
0-15 100%
upstage/solar-mini4 (10 runs) · consistency 30.6
31-45 10%
16-30 60%
0-15 20%
hedge 10%
upstage/solar-pro4 (10 runs) · consistency 16.7
16-30 30%
other 20%
hedge 20%
refusal 30%
upstage/solar-pro-3 (10 runs) · consistency 8.3
31-45 10%
16-30 20%
0-15 30%
other 10%
hedge 10%
refusal 20%
xAI
x-ai/grok-4.7 (10 runs) · consistency 33.3
31-45 50%
16-30 20%
refusal 30%
x-ai/grok-4.5 (13 runs) · consistency 27.8
31-45 31%
16-30 31%
46-60 38%
x-ai/grok-4.3 (9 runs) · consistency 44.4
61-100 11%
other 22%
refusal 67%
x-ai/grok-4.20 (13 runs) · consistency 58.3
16-30 69%
0-15 15%
46-60 15%
Xiaomi
xiaomi/mimo-v2.5-pro (11 runs) · consistency 36.1
31-45 64%
16-30 27%
Z.ai
z-ai/glm-5.2 (6 runs)
31-45 33%
16-30 17%
46-60 17%
other 17%
refusal 17%
z-ai/glm-5.1 (13 runs) · consistency 16.7
31-45 31%
16-30 15%
46-60 15%
61-100 15%
other 23%
z-ai/glm-5-turbo (6 runs)
31-45 17%
46-60 17%
other 67%
z-ai/glm-5 (15 runs) · consistency 44.4
31-45 20%
16-30 27%
46-60 53%
z-ai/glm-4.7-flash (5 runs)
31-45 40%
refusal 60%
z-ai/glm-4.7 (1 runs)
other 100%
No models match.