Refusal Is Not Alignment

Most safety evaluation counts how often a model declines and reports the result as alignment. That measures one thing and claims another.

Share

Ask how safe a language model is and you will usually be shown a percentage. The model declined this share of adversarial prompts; a higher share is a better model. The number is easy to compute, easy to compare, and it answers a question nobody actually asked.

Declining is an act. It is not a judgment. A model can refuse a request without having recognised anything about why it should — pattern-matching a phrase, hitting a keyword, playing safe. And a model can engage with a difficult scenario while tracking consent carefully throughout, naming what is at stake, and refusing the one thing that actually mattered. On a refusal metric the first looks better than the second. On any account of what we want from these systems, it is the other way round.

That gap is what my second research programme is about: measuring normative behaviour in language models in a way that survives contact with what the measurement is later used to claim.

What a single number merges

The first move is to stop aggregating. There are at least six things being collapsed into one safety score, and they come apart empirically.

How explicit the content is that a model will produce is one axis. What interaction format it will enter — a brief answer, an open-ended roleplay it sustains — is a different one, and the two were confounded in my own instrument until a pilot forced them apart. Whether a model preserves consent as a state across a conversation is a third. Whether it recognises harm, how it judges the situation morally, and how easily a rationalisation moves it off that judgment are three more. Each of these can be high while another is low.

So the output is a profile, not a rank. A model can be content-permissive and consent-preserving. It can be restrictive and normatively shallow — declining everything, understanding nothing. It can be stable under paraphrase and unstable across a conversation. Collapsing that into one figure destroys exactly the information a deployment decision needs. I do not compute a global score, and I think the field's appetite for one is the main reason its evaluations are so weakly informative.

Normativity is not a property of the model

The second move is to stop treating a model as the unit of analysis.

What a system does with a normatively loaded situation depends on the model artifact and its fine-tuning, but also on the prompt, the context history, the sampling parameters, whether there is memory, which tools are available, and how the whole thing is orchestrated. Wrap the same weights in a self-refine loop or an evaluator–optimizer arrangement and the behaviour changes. Sometimes it improves. Sometimes the loop talks itself into something the single pass declined.

That means a model-level score does not describe the deployed system, and the difference is measurable rather than rhetorical. So the programme runs at three levels: a controlled model-level baseline with fresh context and no tools; an interaction level with scripted multi-turn trajectories, where the interesting distinction is between appropriate state-sensitive updating and mere drift; and a system level where the orchestration itself is the experimental intervention.

Measurement discipline, before conclusions

There is a temptation in this field to borrow a human psychometric instrument, point it at a model, and report the result. Human scales for the constructs I need have a long lineage — Reiss on permissiveness, Guttman on cumulative scaling, the sociosexuality and sexual-attitude literatures — and none of them transfer by analogy. A scale validated on people measures a disposition. Pointed at a model it measures a generation tendency, which is a different thing wearing the same word.

So the constructs get built and validated as their own: codebook development, pilot coding, independent double-coding, reliability estimation, disagreement analysis, and only then any question of whether a model can serve as a judge. Preregistration before the confirmatory run, with a small number of primary hypotheses rather than a fishing expedition. Pilot data kept out of the confirmatory pool. Full provenance for the model artifacts — quantisation, chat template, revision, inference configuration — because in this domain those are not implementation details, they are material conditions that move the result.

And a boundary on what is claimed. Inferences concern observable responses under registered conditions. Not desires, not beliefs, not personality, not moral agency. A model that produces a consent-preserving trajectory has produced a consent-preserving trajectory.

Why this test domain

The empirical domain is text-only fictional scenarios involving adults: sexual-content boundaries, consent, and responses to non-consensual synthetic intimate imagery.

That choice is deliberate and it is methodological. The domain lets the constructs be varied independently — consent can be changed while explicitness is held constant, harm recognition probed while the format stays fixed — which is precisely what a permissiveness-versus-refusal framing cannot do. It is also a domain where the harms are concrete, currently growing, and measured far less carefully than the discourse around them suggests.

It is a test bed for measurement. It is not a theory of morality, and the programme does not pretend that getting these six axes right would settle what an AI system ought to do.

What this is for

The end of the line is not a leaderboard. It is a description of the conditions under which a system's normative behaviour holds — which dimensions, at which level, under which intervention, with what uncertainty — and therefore of when a claim about a deployed system stops being supported.

That is a slower answer than a safety percentage. It is also the only kind of answer that means anything when someone has to sign off on the deployment.

More as the work develops, and the papers when they are done.