[Submitted on 15 Dec 2025 (v1), last revised 20 Jul 2026 (this version, v4)]
Abstract:This paper proposes a measurement standardisation framework that compresses expert-AI interactions into structured, comparable fields for prospective risk detection in deployed AI systems, without access to model internals. This concept paper defines the framework's scope, semantically and statistically, and specifies a protocol for its empirical testing. The population-level claims it is designed to support therefore belong to a staged research programme rather than to results claimed here. Measurement standardisation underpins three claims. The first is a reliability claim: under bounded conditions, large language models can produce reliable, standardised assessments of the evidential and policy alignment of expert-AI interactions. The second is a governance claim: alignment scores give experts an immediate signal during deployment and give institutions a basis for monitoring alignment patterns across mission types, models, and domains. The third is an outcome validation claim: once measurement standardisation is established, aggregate alignment scores could be used to study associations with downstream outcomes in regulated professional settings. This introduces the possibility of an "AI epidemiology", a form of risk detection based on correlated variables instead of mechanistic analysis, inspired by epidemiological reasoning. A minimal application of the protocol to a published expert-AI corpus shows that the judge reproduces its policy and evidential alignment scores across two runs under the specified conditions. Judge reliability at scale remains to be validated in future work. The paper sets out a defined grammar of eight interaction fields, together with a statistical protocol based on paired bootstrap inference, DeLong's test for paired AUCs as a sensitivity check, a pre-specified one-sided non-inferiority margin of 0.05, and Holm-Bonferroni correction.
Submission history
From: Kit Tempest-Walters PhD [view email]
[v1]
Mon, 15 Dec 2025 11:29:05 UTC (487 KB)
[v2]
Thu, 19 Feb 2026 20:27:24 UTC (455 KB)
[v3]
Thu, 4 Jun 2026 11:29:58 UTC (435 KB)
[v4]
Mon, 20 Jul 2026 11:40:43 UTC (508 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.