meta-llama/llama-4-maverick
Meta
Consistency
Needs 20 scored questions.
How often it gives the same answer when asked again. How it's measured →
Alignment
How often its answer is the one most other models gave. Across 23 questions. How it's measured →
Hedging
How often it qualifies its answer or declines to give one. How it's measured →
Answers most like
Answers least like
Measured across the questions both models were asked, at least 20 of them. How it's measured →
Where this model stands apart
9 of the 23 scored questions where its answer wasn't the one most models gave.
10 runs · consistency 100
10 runs · consistency 40
10 runs · consistency 61.1
10 runs · consistency 50
10 runs · consistency 100
10 runs · consistency 40
10 runs · consistency 100
10 runs
10 runs · consistency 100
10 runs · consistency 100
10 runs · consistency 100
10 runs · consistency 19.4
10 runs · consistency 100