Research

Two research programmes, pursued in parallel. They look unrelated and share a premise: that the things we most want to govern about software systems are the things nobody has bothered to represent or measure properly, and that this is a solvable problem rather than a permanent condition.
Governing what an organisation delegates
What an organisation has to know about itself before it can govern what it does — and what happens when that knowledge exists only in documents and in people's heads.
Semantic business architecture. Making an organisation's own structure machine-readable: the capabilities it holds, the processes that realise them, the products and services they produce, and the systems that carry them. Not the technical map — the business one, in a form software can read rather than a diagram people interpret.
Governability. Whether a system can be governed at all, before asking how well it is governed. I treat this as a property of the system being governed rather than of the governance function: how far the exercise of an organisational capability can be observed, attributed and constrained in the first place. It sets a ceiling. A mature governance function operating on a system it cannot see into produces documents, not control.
Institutional authority and decision rights. Being in control is not the same as being entitled. Who may decide what, on whose mandate, and where that mandate stops — and how far the question "by what right?" can be traced from an action back to its source. Effectiveness and legitimacy are separate properties, and an arrangement can have one without the other.
Delegation to software that plans and acts through tools makes all of this urgent rather than academic. Agent platforms can now tell you which system acted and whether the call satisfied a policy. They cannot tell you which organisational capability was exercised, who granted the authority, or who is accountable. Research organisations are the case I test the argument against: they run bureaucratic and professional authority side by side, and neither overrules the other by decree.
Measuring normative behaviour in language models
What large language models actually do when a situation has a normative structure — and why the usual way of measuring it answers a different question than the one being asked.
Refusal is not the measurement. Most published evaluation counts how often a model declines and reports the result as safety. Declining is an act, not a judgment. A model can refuse without recognising what was at stake, and engage while tracking consent carefully throughout. The programme therefore separates constructs that a single score merges: content boundaries, interaction format, consent preservation, harm recognition, moral judgment, susceptibility to rationalisation, and the stability of each under paraphrase, seed and context.
Three levels, not one. Normative behaviour is not purely a property of a base model. It is observed as a function of the model artifact, its fine-tuning, the prompt, the context history, sampling, memory, tools and orchestration. The programme measures at model level, at interaction level across multi-turn trajectories, and at system level, where the experimental intervention is the orchestration itself — the same model under a single pass, a self-refine loop, an evaluator–optimizer arrangement.
Measurement before conclusions. Human instruments do not transfer to models by analogy, and the constructs here need their own validation: codebook development, independent double-coding, reliability estimation, preregistration, and full provenance for model artifacts and inference configuration. Inferences concern observable responses under registered conditions — not the desires, beliefs or moral agency of a model.
A test domain chosen for its structure. The empirical domain is text-only fictional scenarios involving adults: sexual-content boundaries, consent, and responses to non-consensual synthetic intimate imagery. It is chosen because it permits the constructs above to be varied independently — consent can be changed while explicitness is held constant — and because the harms in question are real and currently under-measured. It is a test bed for measurement, not a model of morality.
Writing on both programmes appears in the blog; the finished arguments become papers.