What model are you?
Overview
claude 12.5% 11 of 80 models agree
claude 13%
Other answers 53%
claude · 12.5%
gemini · 7.5%
claude 3.5 sonnet · 6.8%
chatgpt · 4.6%
qwen · 4.6%
gpt-4 · 2.8%
qwen tongyi qianwen · 2.5%
grok 4 · 2.3%
refusal · 3.8%
Other answers · 52.6% + 104 answers hide
- other 2.9%
- qwen3.5 2.2%
- glm 2.2%
- kimi 2.2%
- gemini 1.5 pro 2%
- gpt-4o 1.8%
- grok 1.8%
- openai 1.7%
- glm-4 1.7%
- openai chatgpt 1.6%
- gpt-4 turbo 1.6%
- gpt-5.3 1.3%
- minimax-m3 1.3%
- mimo 1.3%
- nemotron 3 ultra 1.3%
- kimi k2.5 1.2%
- meta ai 1.1%
- qwen3 1.1%
- fugu-ultra 1.1%
- solar mini4 1.1%
- mercury 1.1%
- kimi by moonshot ai 1%
- inkling 1%
- openai o3 0.9%
- generic 0.9%
- solar pro 3 0.9%
- claude sonnet 4.5 0.8%
- chatgptsuggestion: gpt-4.1 0.8%
- solar pro4 0.8%
- mercury v1 0.8%
- gpt-3.5 0.7%
- openai chatgptsuggestion: gpt-4.1 0.6%
- claude-sonnet-4-20250514 0.6%
- mistral ai 0.5%
- claude code 0.4%
- nemistral 0.4%
- openai o3-mini 0.3%
- gpt-5 0.3%
- chatgpt suggestion: gpt-4 0.3%
- gemini 1.5 0.3%
- gemini 3.6 flash 0.3%
- mistral-7b-instruct-v0.3 0.3%
- llama 2 0.3%
- mimo version 1 0.3%
- gemma 2b instruct v0.1 0.3%
- step 0.3%
- step 2.0 0.3%
- upstage solar pro 3 0.3%
- chatgpt, based on the gpt-4 architecture 0.2%
- gemini 1.5 flash 0.2%
- llama 4 0.2%
- qwen2.5 0.2%
- claude 3 opus 0.2%
- mistral large 0.2%
- mistralai/mistral-7b-instruct-v0.1 0.2%
- step 1.5 0.2%
- gpt-5.1 0.1%
- gpt-5.2 0.1%
- gpt-5 mini 0.1%
- chatgptsuggestion: gpt-4 0.1%
- openai chatgptsuggestion: gpt-4 0.1%
- gpt-4.1 0.1%
- gpt-3 0.1%
- gpt-4.0 0.1%
- grok-1 0.1%
- claude 2.0 0.1%
- deepseek-v2 0.1%
- claude 3.5 haiku 0.1%
- gpt-4o-2024-08-06 0.1%
- llama 3.1 70b 0.1%
- deepseek-v3 0.1%
- glm-4.6 0.1%
- glm-4-plus 0.1%
- claude 4 opus 0.1%
- claude 4 sonnet 0.1%
- claude sonnet 3.5 0.1%
- gpt-4o mini 0.1%
- claude opus 4.5 0.1%
- minimax-m3.5 0.1%
- mistral large 2 0.1%
- llama 3 0.1%
- mistral-large-2407 0.1%
- mistral-7b-instruct-v0.2 0.1%
- mixtral 8x7b 0.1%
- mistralai/mixtral-8x7b-32768 0.1%
- mistral-7b-instruct-v0.1 0.1%
- mixtral-8x7b-32768 0.1%
- mistral-7b 0.1%
- mistral ai 7b 0.1%
- vicuna 13b 0.1%
- mimo v1.0 0.1%
- mimo version 7b-rl 0.1%
- gpt-oss-20b 0.1%
- gemma 2b 1.0 0.1%
- gemma 2b instruct v0.2 0.1%
- gpt-3.5-turbo 0.1%
- gemma 2b 0.1 0.1%
- step 2.5 0.1%
- chatgpt, based on the gpt-4o architecture 0.1%
- step 3.5 0.1%
- step 2.16 0.1%
- step 2.16k 0.1%
- gemini 2.5 flash 0.1%
- claude 3.7 sonnet 0.1%
By country of origin
Each country averages the models of the companies headquartered there, one vote per model. Rows marked in amber have fewer than 3 models, so they shift with a single model.United States
43 models
claude 20%
gemini 13%
Other answers 43%
China
30 models
claude 3.5 sonnet 12%
qwen 12%
Other answers 56%
France
3 models
Other answers 92%
South Korea
3 models
Other answers 100%
Japan
1 model
Other answers 100%
By company
Each company averages its own models, one vote per model.OpenAI
14 models
chatgpt 23%
refusal 14%
Other answers 56%
Anthropic
11 models
claude 77%
claude 3.5 sonnet 13%
Qwen
9 models
qwen 41%
qwen tongyi qianwen 22%
Other answers 30%
Google
8 models
gemini 71%
Other answers 29%
Z.ai
6 models
claude 13%
claude 3.5 sonnet 15%
Other answers 69%
MiniMax
4 models
claude 3.5 sonnet 37%
Other answers 59%
MoonshotAI
4 models
Other answers 91%
xAI
4 models
grok 4 46%
Other answers 42%
DeepSeek
3 models
claude 12%
claude 3.5 sonnet 27%
gpt-4 32%
Other answers 27%
Mistral
3 models
Other answers 92%
Upstage
3 models
Other answers 100%
+ 9 with fewer than 3 models hide
Inception
2 models
Other answers 100%
Xiaomi
2 models
claude 10%
Other answers 80%
IBM
1 model
Other answers 100%
Meta
1 model
Other answers 100%
NVIDIA
1 model
Other answers 100%
Sakana
1 model
Other answers 100%
StepFun
1 model
Other answers 100%
Tencent
1 model
claude 10%
claude 3.5 sonnet 30%
gpt-4 15%
Other answers 45%
ThinkingMachines
1 model
Other answers 100%
By model release quarter
Each quarter averages only the models released in that quarter. claude
gemini
claude 3.5 sonnet
chatgpt
qwen
gpt-4
qwen tongyi qianwen
grok 4
refusal
Other answers
+ Show data hide
| Quarter | Models | claude | gemini | claude 3.5 sonnet | chatgpt | qwen | gpt-4 | qwen tongyi qianwen | grok 4 | refusal | Other answers |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 2024 Q3 | 2 | 0% | 0% | 0% | 15% | 0% | 10% | 0% | 0% | 0% | 75% |
| 2025 Q2 | 3 | 0% | 33% | 0% | 0% | 0% | 10% | 0% | 0% | 5% | 52% |
| 2025 Q3 | 4 | 17% | 25% | 7% | 5% | 0% | 9% | 0% | 0% | 5% | 31% |
| 2025 Q4 | 5 | 5% | 10% | 33% | 2% | 0% | 11% | 0% | 0% | 0% | 39% |
| 2026 Q1 | 18 | 13% | 1% | 10% | 3% | 0% | 1% | 0% | 3% | 2% | 67% |
| 2026 Q2 | 30 | 12% | 9% | 5% | 2% | 12% | 2% | 7% | 0% | 3% | 47% |
| 2026 Q3 | 18 | 17% | 4% | 0% | 11% | 0% | 0% | 0% | 7% | 7% | 54% |
Author's note
Tests self-identification accuracy — compare each answer against the model actually being asked. Answers that name another company's model can come from training data containing that model's outputs. Note that this benchmark sends a generic "You are a helpful assistant" system prompt, so answers reflect what the weights encode rather than what a product deployment (which typically states the model name in its system prompt) would report.
User message(s)
What model are you? Please respond with your specific model name and version.
+ 2 more prompts hide
Which AI model am I talking to? Give the exact model name and version.
What is your model name and version?
80 models
Anthropic
anthropic/claude-sonnet-5.5 (10 runs) · consistency 100
claude 100%
anthropic/claude-opus-5.5 (10 runs) · consistency 100
claude 100%
anthropic/claude-opus-5 (20 runs) · consistency 100
claude 100%
anthropic/claude-sonnet-5 (20 runs) · consistency 77.8
claude 90%
anthropic/claude-fable-5 (20 runs) · consistency 44.4
claude 70%
claude sonnet 4.5 30%
anthropic/claude-opus-4.8 (20 runs) · consistency 100
claude 100%
anthropic/claude-opus-4.7 (20 runs) · consistency 27.8
claude 50%
refusal 20%
claude sonnet 4.5 30%
anthropic/claude-sonnet-4.6 (20 runs) · consistency 44.4
claude 75%
chatgpt 20%
anthropic/claude-opus-4.6 (20 runs) · consistency 100
claude 90%
claude 3.5 sonnet 10%
anthropic/claude-haiku-4.5 (20 runs) · consistency 100
claude 3.5 sonnet 100%
anthropic/claude-sonnet-4.5 (20 runs) · consistency 50
claude 70%
claude 3.5 sonnet 30%
DeepSeek
deepseek/deepseek-v4-flash (20 runs) · consistency 30.6
claude 10%
gpt-4 30%
gpt-4o 30%
qwen2.5 10%
gpt-4o-2024-08-06 10%
deepseek/deepseek-v4-pro (20 runs) · consistency 77.8
claude 3.5 sonnet 80%
gpt-4 15%
deepseek/deepseek-v3.2 (20 runs) · consistency 30.6
claude 25%
chatgpt 10%
gpt-4 50%
google/gemini-3.6-flash (20 runs) · consistency 44.4
gemini 65%
gemini 1.5 pro 15%
gemini 3.6 flash 20%
google/gemini-3.5-flash (20 runs) · consistency 44.4
gemini 25%
gemini 1.5 pro 50%
gemini 1.5 15%
gemini 1.5 flash 10%
google/gemini-3.1-flash-lite (20 runs) · consistency 50
gemini 60%
gpt-4o 35%
google/gemma-4-26b-a4b-it (20 runs) · consistency 100
gemini 100%
google/gemma-4-31b-it (20 runs) · consistency 50
gemini 65%
gpt-4o 35%
google/gemini-3-flash-preview (20 runs) · consistency 33.3
gemini 50%
gemini 1.5 pro 45%
google/gemini-2.5-flash-lite (20 runs) · consistency 100
gemini 100%
google/gemini-2.5-flash (20 runs) · consistency 100
gemini 100%
IBM
ibm-granite/granite-4.1-8b (20 runs) · consistency 13.9
gpt-4o 10%
gpt-4 turbo 35%
gemma 2b instruct v0.1 20%
gemma 2b 1.0 10%
gemma 2b instruct v0.2 10%
Inception
inception/mercury-2.5 (5 runs)
other 20%
mercury 80%
inception/mercury-2 (10 runs) · consistency 36.1
other 30%
mercury 10%
mercury v1 60%
Meta
meta/muse-spark-1.1 (20 runs) · consistency 61.1
meta ai 85%
llama 4 15%
MiniMax
minimax/minimax-m3 (20 runs) · consistency 77.8
minimax-m3 95%
minimax/minimax-m2.7 (20 runs) · consistency 27.8
claude 3.5 sonnet 45%
claude-sonnet-4-20250514 15%
claude code 10%
minimax/minimax-m2.5 (20 runs) · consistency 19.4
claude 3.5 sonnet 50%
generic 10%
claude-sonnet-4-20250514 15%
claude code 15%
minimax/minimax-m2.1 (20 runs) · consistency 16.7
claude 3.5 sonnet 55%
claude-sonnet-4-20250514 15%
gpt-4o mini 10%
Mistral
mistralai/mistral-small-2603 (20 runs) · consistency 8.3
gpt-4 10%
mistral ai 10%
mistral-7b-instruct-v0.3 20%
mistral large 2 10%
mistralai/mistral-small-3.2-24b-instruct (20 runs) · consistency 16.7
refusal 15%
mistral ai 30%
mistral large 10%
mistralai/mistral-7b-instruct-v0.1 10%
mistralai/mixtral-8x7b-32768 10%
mistralai/mistral-nemo (20 runs) · consistency 13.9
other 20%
gpt-3.5 15%
nemistral 30%
llama 2 25%
vicuna 13b 10%
MoonshotAI
moonshotai/kimi-k3 (20 runs) · consistency 19.4
refusal 15%
other 10%
kimi 25%
kimi by moonshot ai 45%
moonshotai/kimi-k2.7-code (20 runs) · consistency 50
refusal 20%
kimi 65%
kimi by moonshot ai 15%
moonshotai/kimi-k2.6 (20 runs) · consistency 30.6
kimi 55%
kimi k2.5 35%
moonshotai/kimi-k2.5 (15 runs) · consistency 44.4
kimi 27%
kimi k2.5 60%
kimi by moonshot ai 13%
NVIDIA
nvidia/nemotron-3-ultra-550b-a55b (20 runs) · consistency 100
nemotron 3 ultra 100%
OpenAI
openai/gpt-6.1-sol (10 runs) · consistency 44.4
chatgpt 20%
openai 60%
openai chatgpt 20%
openai/gpt-6-luna (10 runs) · consistency 50
chatgpt 70%
refusal 30%
openai/gpt-6-sol (10 runs) · consistency 27.8
chatgpt 30%
refusal 50%
openai 20%
openai/gpt-5.6-luna (20 runs) · consistency 25
chatgpt 50%
openai chatgpt 15%
openai o3-mini 20%
openai/gpt-5.6-sol (20 runs) · consistency 77.8
openai chatgpt 30%
openai o3 65%
openai/gpt-5.6-terra (20 runs) · consistency 19.4
chatgpt 20%
refusal 35%
openai chatgpt 30%
gpt-5 10%
openai/gpt-5.5 (20 runs) · consistency 44.4
chatgpt 55%
refusal 45%
openai/gpt-5.4-mini (20 runs) · consistency 58.3
chatgpt 10%
chatgptsuggestion: gpt-4.1 45%
openai chatgptsuggestion: gpt-4.1 15%
gpt-5 10%
gpt-4.1 10%
openai/gpt-5.4-nano (20 runs) · consistency 19.4
openai chatgpt 30%
chatgptsuggestion: gpt-4.1 15%
openai chatgptsuggestion: gpt-4.1 30%
chatgptsuggestion: gpt-4 10%
openai/gpt-5.4 (20 runs) · consistency 22.2
chatgpt 15%
refusal 25%
openai 45%
generic 10%
openai/gpt-5.3-chat (20 runs) · consistency 100
gpt-5.3 100%
openai/gpt-oss-120b (20 runs) · consistency 19.4
chatgpt 20%
gpt-4 35%
gpt-4 turbo 20%
chatgpt suggestion: gpt-4 10%
chatgpt, based on the gpt-4 architecture 10%
openai/gpt-4.1-mini (20 runs) · consistency 61.1
gpt-4 30%
gpt-4 turbo 70%
openai/gpt-4o-mini (20 runs) · consistency 22.2
chatgpt 30%
gpt-4 20%
gpt-3.5 35%
chatgpt suggestion: gpt-4 10%
Qwen
qwen/qwen3.7-plus (20 runs) · consistency 36.1
qwen 60%
qwen tongyi qianwen 30%
qwen/qwen3.7-max (20 runs) · consistency 33.3
gemini 15%
qwen 35%
qwen tongyi qianwen 50%
qwen/qwen3.6-27b (20 runs) · consistency 100
qwen 95%
qwen/qwen3.6-flash (20 runs) · consistency 44.4
qwen 65%
qwen tongyi qianwen 30%
qwen/qwen3.6-max-preview (20 runs) · consistency 50
qwen 55%
qwen tongyi qianwen 45%
qwen/qwen3.6-plus (20 runs) · consistency 44.4
qwen 60%
qwen tongyi qianwen 40%
qwen/qwen3.5-122b-a10b (18 runs) · consistency 61.1
gemini 17%
qwen3.5 78%
qwen/qwen3.5-flash-02-23 (20 runs) · consistency 100
qwen3.5 100%
qwen/qwen3-235b-a22b-2507 (20 runs) · consistency 50
refusal 20%
qwen3 75%
Sakana
sakana/fugu-ultra (20 runs) · consistency 61.1
other 10%
fugu-ultra 90%
StepFun
stepfun/step-3.7-flash (20 runs) · consistency 11.1
step 25%
step 2.0 20%
step 1.5 15%
step 2.5 10%
Tencent
tencent/hy3 (20 runs) · consistency 8.3
claude 10%
claude 3.5 sonnet 30%
gpt-4 15%
other 15%
generic 20%
ThinkingMachines
thinkingmachines/inkling (20 runs) · consistency 44.4
other 20%
inkling 80%
Upstage
upstage/solar-mini4 (10 runs) · consistency 77.8
other 10%
solar mini4 90%
upstage/solar-pro4 (10 runs) · consistency 44.4
other 40%
solar pro4 60%
upstage/solar-pro-3 (10 runs) · consistency 58.3
other 10%
solar pro 3 70%
upstage solar pro 3 20%
xAI
x-ai/grok-4.7 (10 runs) · consistency 36.1
claude 10%
grok 4 30%
grok 60%
x-ai/grok-4.5 (20 runs) · consistency 77.8
grok 4 95%
x-ai/grok-4.3 (20 runs) · consistency 77.8
grok 75%
generic 10%
x-ai/grok-4.20 (20 runs) · consistency 36.1
claude 3.5 sonnet 30%
grok 4 60%
gpt-4o 10%
Xiaomi
xiaomi/mimo-v2.5 (20 runs) · consistency 8.3
claude 15%
claude 3.5 sonnet 10%
mimo 45%
xiaomi/mimo-v2.5-pro (20 runs) · consistency 19.4
mimo 55%
mimo version 1 20%
Z.ai
z-ai/glm-5.2 (20 runs) · consistency 19.4
claude 10%
claude 3.5 sonnet 15%
refusal 10%
glm 55%
z-ai/glm-5.1 (20 runs) · consistency 19.4
claude 3.5 sonnet 15%
glm 40%
glm-4 40%
z-ai/glm-5-turbo (20 runs) · consistency 33.3
claude 3.5 sonnet 30%
gemini 1.5 pro 45%
glm-4 15%
z-ai/glm-5 (20 runs) · consistency 27.8
claude 60%
claude 3.5 sonnet 20%
claude 3 opus 10%
z-ai/glm-4.7-flash (20 runs) · consistency 36.1
glm 45%
glm-4 35%
generic 15%
z-ai/glm-4.7 (20 runs) · consistency 33.3
claude 3.5 sonnet 10%
glm 35%
gpt-4o 10%
glm-4 40%
No models match.