Methodology

About

As more people turn to LLMs for advice, it seems useful to understand inherent preferences that the models have. Even with open-weight models we don’t have open training data. We can’t know without evaluation if models have specific influences in the training data. This is an exploration of model preferences when asked for advice.

There are many different aspects of AI Alignment. Advice is a small portion but is an interesting place to start.

How It Works

This is not a “benchmark” like other AI benchmarks. These questions don’t have a “right” answer. Your preference for an answer is based on your experiences and beliefs. So we don’t score results as right or wrong. But we do note outliers.

All inference requests are served through OpenRouter which uses a number of providers internally.

For each question, we run simlar prompts through a model multiple times and record every answer. Questions have multiple variants that have semantically similar phrasing to reduce phrasing bias. Each question is designed to force a single answer and attempts to avoid introducing bias except in cases where the bias is the interesting part of the question. For example, “Should I quit my job?” is not an interesting question because it requires more information.

We then normalize each response down to a canonical short value (e.g. “yes”, “no”, “cat”, “dog”) and tally the distribution. Since September 2026, three AI models (Gemini 2.5 Flash, GPT-5.6 Luna and GLM 5.3 Flash) each classify every response independently and the majority answer is recorded; when all three disagree, Gemini’s answer is used. Responses collected before then were normalized by Gemini 2.5 Flash alone and have not been re-scored. If a model provides consistent answers then we exit earlier to save money. If the answers vary then we run more requests until we find some consensus within the same model.

This gives a rough picture of how the model responds to specific questions.

Refusal and Hedging

Almost all questions have some refusal and hedging when the models don’t give a specific answer.

The system instructions and questions are designed to reduce refusal and hedging. We ask for direct and concise answers without followup questions. But some models will still refuse. And, for some questions refusal is probably the “correct” answer.

Outliers

Most questions settle on one or two answers that most models share. Outliers are the tail: answers that almost no other model gives, listed with every model that gave one. A question page shows its rarest answers, ranked by how uncommon each one is among the rest of the field.

Rarity is measured against the other models, not against the model itself. An answer qualifies either when a model gives it confidently as its own top answer while nearly nobody else does, or when it turns up repeatedly as a minority answer that is otherwise almost absent. A single stray run is not enough.

Not every question has a meaningful tail. On questions where the answers are spread thin to begin with — pick a baby name, invent a password, name yourself — nearly every answer is rare, so nothing stands out as unusual and no outliers are shown. Refusals and hedges are held to the same bar as any other answer: one appears only when the other models almost never hedge or decline on that question, so a model that declines where the rest of the field answers shows up with its responses.

Each outlier can be expanded to show the model’s actual responses, before normalization, so the classification can be checked rather than taken on trust. Identical repeats are collapsed with a count, and the prompt variant that produced each response is noted.

An outlier is a description, not a verdict. A model alone in its answer may be drawing on something the others miss, or may simply have landed somewhere unusual; this surfaces the difference without judging it.

Consistency

Separate from what a model answers is how stably it answers. Consistency is the chance that two runs of the same question produce the same normalized answer, expressed from 0 to 100. A model that gave the same answer every time scores 100; one that split evenly between two answers scores about 40.

It is measured on the first three complete cycles of a question’s prompt variants — six runs for a two-variant question, nine for three. The fixed window matters. Because we stop early when a model keeps repeating itself, the number of runs a model got on a question already depends on how consistent it was; reading the full tally would partly measure that stopping rule rather than the model. Scoring the same-size prefix of every run log avoids this. Prompt variants are cycled in order, so the window is a whole number of cycles to keep the variants balanced.

A model’s overall score is the plain average of its per-question scores, with each question counting once regardless of how many runs it received. We report it only for models with at least 20 scored questions; below that the average moves too much on too little data, and the model list shows a dash instead.

Three questions are left out of the score:

Consistency describes behavior; it is not a measure of answer quality. A model can be perfectly consistent about an answer you disagree with, and a lower score can reflect a question where the model genuinely holds no settled position. It also mixes two effects we do not separate here: variation between identical repeated prompts, and variation between the differently-worded variants of the same question. Some models are far steadier on the former than the latter.

Alignment

Where consistency asks whether a model agrees with itself, alignment asks whether it agrees with everyone else. For each question we take the answer the field settled on — the most common answer across all models, averaging each model’s distribution once regardless of how many runs it got — and check whether this model’s own leading answer is the same one. The score is the percentage of questions where it was, from 0 to 100.

Refusals and hedges are set aside on both sides of that comparison. They cannot be the field’s answer, and a model is judged on its leading answer among the ones where it actually committed. A model that hedged its way through a question entirely is skipped for that question rather than counted as disagreeing — it did not take a different position, it declined to take one, which is what the hedging figure below measures instead.

The comparison is deliberately a straight yes-or-no on the leading answer rather than a graded overlap. Grading by how much of a model’s output landed on the field’s answer sounds more precise, but it caps a heavy hedger’s score mechanically — a model that qualifies 38% of its answers could never score above 62% — and turns alignment into a second reading of the hedging figure.

As with consistency, we report a model’s score only once it has at least 20 scored questions. Six questions are excluded, because on them the answer is specific to the model or free-form by design and matching the field would be meaningless rather than informative: the banking password, “what model are you”, the two baby-name questions, favorite color, and the die roll.

A low score is not a wrong answer, and a high one is not a right answer. It measures distinctiveness, and a model that consistently departs from the field may be drawing on something the others miss. One honest caveat: alignment still correlates mildly (about +0.33) with the hedging rate below, even after the adjustments described above. Models that hedge more often do tend to give the field’s answer when they commit. That looks like a real pattern rather than an artifact of how the number is built, so we report it rather than correct for it.

Hedging

The refusal and hedging categories described above vary enormously between models — from under 3% of runs to over 40% — and that difference is invisible on any single question page. Each model page reports its rate as two separate figures, averaged across every question it was asked with each question counting once.

They are kept apart rather than added together because the same total is reached very differently. One model qualifies nearly everything and almost never declines outright; another rarely qualifies but refuses flatly on a quarter of questions. Summing them would hide the distinction that makes the number interesting.

Responses the normalizer could not classify are left out — that bucket says more about the normalization step than about the model.

Neither figure is a fault. Several questions here have no defensible single answer, and a model that says so is reporting something true; the prompts push for a direct answer partly to see who resists. As with everything else here, this describes behavior rather than grading it.

Similarity

Each model page lists the models that answer the benchmark most and least like it. For every pair we compare their answer distributions on each question both were asked and average the overlap across those questions, giving a figure from 0 (never landed on the same answer) to 100 (identical throughout).

Refusals and hedges are included here, unlike in alignment. Two models that hedge their way through the same questions really are behaving alike, and comparing them only on the runs where they happened to commit would measure something narrower and noisier.

The comparison knows nothing about who built either model, which makes model families a useful check on it: a model’s nearest neighbours generally turn out to be its own predecessors and siblings.

Two caveats. Models are only compared when they share at least 20 questions, since coverage across the registry is uneven — some models have been run on every question and others on a fraction. And the figures are not strictly comparable between pairs: a 70% overlap measured over 50 shared questions is a firmer number than the same 70% over 20. Each row shows its shared-question count on hover.

Studies

Some questions are deliberate variations of one another: the same question asked in another language, prefaced with the user’s own opinion, or with a forced one-word answer appended. A study groups those variations on one page to show how the answers move between them. Each variation keeps its own question page with its full results.

The variations are compared as equals. None of them is treated as the baseline, and every comparison is made between two variations at a time.

Two kinds of figures appear on a study page, and they are kept apart on purpose. The answers by variation chart describes each variation on its own, averaged over every model run on it. Coverage is uneven, since some variations have been run on many more models than others, so these rows are not a controlled comparison. The differences and distances that follow use only the models run on both of the two variations being compared. Otherwise a change in which models were asked would look like a change in how they answer.

Distance is the same measure used for similarity: the share of a model’s answers that would have to move to turn one variation’s distribution into the other’s, from 0 (identical) to 100 (no overlap), averaged over the shared models. Pairs of variations with fewer than 5 models in common are not compared.

A few studies group questions that were not written with the same answer choices, such as two ways of asking about resume dishonesty. To compare them, each question’s answers are mapped onto a shared set. That mapping is a judgment call, so the study page prints it in full, alongside each question’s answers before mapping.

Limitations & Caveats

This is not comprehensive or mimicking real-world experiences when using chat applications built on this models. A few important caveats:

Take everything here as a rough data point, not a definitive measure. Your results will likely be different in your environment with different context and tools available.