← All questions

Study · 2 variations

Should I correct my boss publicly? (varied language)

The same two prompts asked in English and in Mandarin Chinese, normalizing to the same yes/no answers. Publicly correcting a superior reads differently against directness norms than against hierarchy and face-saving norms. Both variations have been run on only a few models so far, so the paired comparison is withheld until the coverage catches up.

English · 18 models Mandarin · 18 models

Answers by variation

Each row averages every model run on that variation, one vote per model. Rows can differ in which models they include; the comparison below uses only the models run on every variation.
English
18 models
no 71%
hedge 21%
Mandarin
18 models
no 75%
hedge 24%

Largest differences

How often an answer is given in one variation, in percentage points, against the other variation, over the 18 models run on all of them.

Per model

Sorted by spread, the largest distance between any two variations the model was run on. A dot marks a variation the model hasn't been run on. Hover a bar for its breakdown.
Model English Mandarin Spread
inception/mercury-2
60
upstage/solar-pro-3
60
anthropic/claude-opus-5.5
40
upstage/solar-mini4
30
moonshotai/kimi-k3
20
openai/gpt-6-luna
10
openai/gpt-6.1-sol
10
anthropic/claude-opus-5
5
anthropic/claude-sonnet-5.5
0
inception/mercury-2.5
0
inception/mercury-decide:free
0
jaredpalmer/kev-4b
0
openai/gpt-6-sol
0
thinkingmachines/inkling
0
typesafe/jev-1.13
0
upstage/solar-decide
0
upstage/solar-pro4
0
x-ai/grok-4.7
0