A research prototype, not an operational library
The implementation is complete relative to its specification and is not operationally validated. The 171 tests, 43 killed mutants and 37 regression cases establish conformance and resistance to the faults tested. They establish nothing about field usability, integration across vendor systems, the truthfulness of supplied provenance, the completeness of construction-boundary capture, or operational safety.
One case fixes the boundary better than any caveat. A spreadsheet applied a fixed formula to a figure a model had produced and presented the result as a calculation. The calculation was correct, nothing recorded that the input was generated, and the demand excluding generated ancestry found none and admitted. The kernel behaved consistently on the record it received; the record represented everything the construction process supplied; and the resulting admission was unwarranted.
Origin conservation conserves what crossed the construction boundary and nothing else. An engineer who reads an estimate on a screen, performs the arithmetic by hand and enters the result defeats the mechanism entirely, and no record anywhere shows it. For a library intended to receive claims from existing plant tooling, the strength of the discipline is bounded by the coverage and truthfulness of that boundary, and neither is a property of the formal model. Any deployment claim would have to be made about the integration, not about the artefact.
Further work
Exposing machinery that exists and cannot be reached. Four corroboration cases failed because the model already holds what they need: origin disjointness is computed, an independence attestation carrying an author and a basis is required before corroboration, and a verification type carrying its own warrant exists. None is reachable through the demand interface. This is latent capability rather than operational representability, and it remains a v1 failure because a user cannot ask for it. It is the lowest-complexity response identified, though whether exposure yields exact translations can only be settled by translating against a revised language, and deciding what a demand may ask about a basis is a semantic commitment with security consequences of its own.
Deciding whether non-entailment needs an inference-policy layer. The present language asks whether a record possesses required properties. A relation asserting that one fact does not license another asks a different kind of question, and adding a general operator would change the artefact's semantic category. The classified failure set suggests a narrow declarative clause naming a condition and the requirement it fails to satisfy, checkable by comparison rather than inference, would answer a substantial part without importing an undeclared logic. That is a hypothesis for v2 rather than a conclusion of v1.
Construction-boundary integration. The observability failure is not a language problem and no operator could address it. What it needs is coverage of the points at which claims enter, and attribution of what crosses them, which is an integration and governance question. This is where the artefact's practical viability will be decided.
Usability evaluation with operational-technology engineers. No engineer authored a demand, constructed a record or read a verdict in this work. Whether the distinctions the model requires can be maintained by the people who would have to maintain them is untested, and the construction matrix for the first adapter suggests it is the harder question: exactly one field of a historian record was directly observed, and six depended on configuration a plant may not have, may have wrong, or may have with nobody willing to attest to it.
Probabilistic systems in reviewing and decision roles
This work assumes at one placement that a refusal shown at a human interface may inform a person's judgement, while acknowledging that disclosure cannot itself block an act. That assumption does not survive unchanged when a probabilistic system occupies the reviewing role. A model receiving a refusal may ignore it, reinterpret it, or reason around it, so a refusal placed before a model is disclosure and not a gate.
A probabilistic reviewer's judgement is itself a generated claim. If that judgement licenses an action, its basis should travel and be adjudicated for that action like any other claim, rather than arriving as a privileged flag that some check was passed. Invoking a second model does not by itself establish verification or independence: two reviewers may share weights, training data, retrieval sources, instrumentation or orchestration, and the model's own position already holds that disjoint identifiers remove a known-dependence veto while granting nothing.
The point generalises beyond any particular deployment. Automation can remove a person from an operational decision loop without settling who authors the standards that govern it, how a reviewer's independence is established, or where responsibility falls when the loop is wrong. That a system performs a task competently does not establish that the authority, independence and accountability attached to the role have transferred with it. This is the failure this work names, arriving at the level of roles rather than of claims: competence at producing an answer treated as authority to act on it.
Nothing here requires a person at every boundary. Humans may leave the execution loop while remaining the authors of the constraints, the owners of the assets, the regulators of permissible conduct and the parties exposed to the consequences. What the discipline needs is not a human at the receiving end but an attributable demand, authored by something accountable, enforced by a mechanism outside the component being governed.
Future work should therefore separate five roles that a single interface can silently merge: the authority that determines what an act requires of its evidence, the producer of a claim, the reviewer that produces a claim about another claim, the adjudicator that mechanically compares basis against demand, and the enforcement point that makes an inadmissible act unavailable. Where one component occupies several, the overlap should be visible rather than dissolved. The adjudicator in particular belongs outside the probabilistic component whose proposals it governs, since a reviewer that can reinterpret its own refusal is not governed by it.
That programme needs its own frozen corpus, and the shape of it is already clear: a model verifying its own output, a second invocation of the same model verifying the first, two reviewers sharing a retrieval source or upstream instrument, a deterministic parser laundering a generated review, a component authoring or relaxing the demand governing its own proposal, and a case in which no person is available but a pre-authorised deterministic policy remains enforceable. Those cases would test whether automation, independence, authority and enforcement can be told apart mechanically.
This extends the work's central failure from claims crossing a boundary to authority transferring implicitly when an automated system takes over a role formerly held by a person. It is a direction the evaluation makes more urgent rather than less: machine-to-machine handovers are more numerous, faster, and afford fewer opportunities for anyone to notice that a claim has acquired standing it never earned. It is not something this version does.
A second version, properly separated. Any revised artefact needs its own formal account, specification, mutation operators, translation, freeze and independently authored holdout corpus. The thirty-nine cases here may show that identified losses are now represented, which is regression evidence. They cannot confirm a version they shaped.
What this work establishes
Action-relative admissibility can be represented and mechanically evaluated for a bounded class of claims. The evaluation also shows that the v1 demand language cannot express most requirements in the frozen corpus, particularly requirements stating that an available fact does not itself warrant a conclusion.
The contribution is therefore not a validated solution for interfacing machine-generated claims with operational technology. It is a formally disciplined prototype, an executable account of its own boundaries, and empirical evidence that property-based demands alone are insufficient for many externally grounded operational warrants. That the boundaries are stated in measurements rather than in caveats is what the sealed protocol was for.