Everything to this point has arranged the evaluation so that it cannot be gamed by its author. The typing was read off the review, the anti-circularity commitment was inherited from the paradigm, the composition criteria were quoted from other people's specifications, and the two faces of failure were declared in advance. All of that establishes that the right thing was built and defines what failing would mean. None of it answers the question the work is actually falsifiable on, which is whether the discipline discriminates. This section is about that question, and it is the one part of the methodology that was genuinely unsolved rather than merely unwritten.
The falsifiable claim, from the problem statement, is that the discipline rejects False Determinism in realistic scenarios without rejecting so much legitimate practice as to be unusable. That is a claim about discrimination, and discrimination is only meaningful against a ground truth. Something has to say which scenarios ought to be refused and which ought to pass, or there is nothing for the discipline's verdicts to be right or wrong about. The difficulty, and it is the same difficulty the whole thesis turns on, is that the obvious sources of that ground truth put a human back exactly where the work says none should stand.
Why the obvious answers fail
Three candidates present themselves, and each fails in a way worth stating, because the route that works is shaped by their failures.
Pre-registration alone is insufficient, and the reason is visible in what the practice was designed to fix. The flexibility available in analytical choices, the researcher degrees of freedom, combined with selective reporting of the most favourable result, raises the rate of false positive and overoptimistic findings (Mandl et al., 2024), and public pre-registration is the reform aimed squarely at it, alongside data sharing and stricter reporting policies (Fletcher, 2024). Fixing the scenarios and their expected verdicts before the rules harden solves the timing half of the circularity here in the same way, since rules settled after the cases cannot be quietly fitted to them. But every one of those safeguards governs when a choice is fixed, not who fixes it, and that is the half this work cannot borrow. It does nothing about authorship. The verdicts are still the designer's honest predictions, and honest is not the same as independent. Pre-registration is necessary and it is not enough.
Expert labelling was rejected here on the ground that a vouch moved is not a vouch removed, and that rejection was too strong. The discipline does not abolish normative judgement and never claimed to: it requires an accountable person to author what an act demands of its evidence, and its contribution is that the judgement is made once, in advance, attributably, and then applied mechanically rather than improvised at each handover. An expert who establishes an evaluation criterion in advance is doing that same ex ante work. What must be refused is not human judgement in the evaluation but the designer silently supplying it, and those are different objections.
Incident reports cover only half the question. They document what went wrong, and in this work they are already bounded to illustration rather than evidence. Even taken at their strongest they give only the false-positive side, the cases that failed, and the harder half of discrimination is the legitimate practice that must not be refused. Failures are written up where normality is not: there is no incident report titled the plant ran correctly on Tuesday, and so no natural corpus of the cases a good discipline must let through.
The route: a field that already evaluates refusal
The way through is that this problem is not new, and the field that solved it is one the review already surveyed and read in full. Selective prediction and learning to defer build systems whose whole job is sometimes to decline, and they face the identical question: how do you evaluate a mechanism whose output is sometimes a refusal? Selective prediction gives a network an integrated reject option and evaluates it by trading coverage against the error rate on what it does answer (Geifman and El-Yaniv, 2019); learning to defer routes a case to a downstream expert on that expert's expected accuracy rather than on the model's own uncertainty (Mozannar and Sontag, 2020). Both answer the evaluation question without labelling any individual case as one that ought to have been refused, and the instrument is the coverage and risk relationship.
What transfers from that field is a set of concepts and not its instrument. Abstention, coverage and risk are the useful borrowings. The curve is not, and it was claimed to be at an earlier stage of this work.
A curve in selective prediction comes from varying a threshold on a scalar the model emits, which orders the cases and yields a family of decisions. This discipline has no such scalar by construction: the review ruled out confidence as the wrong category, the combination is conjunctive with no exchange rate between axes, and there is no global strictness parameter whose variation would order anything. Relaxing a freshness tolerance, admitting a generated ancestry and widening a required scope are three different normative changes, not three points on one scale. Worse, a demand is translated from a reference warrant authored outside this work, so varying the demand to manufacture operating points would change the standard the artefact is being judged against rather than the artefact's own strictness. That is a different task, not a different threshold.
Coverage also splits in two here, because abstention and refusal are different things in this discipline where selective prediction has only one. Decision coverage is the proportion of cases reaching a substantive verdict, admit or refuse, and it measures whether the available records and demands permit adjudication at all. Action coverage is the proportion admitted, and it measures how much practice the discipline permits. Undetermined reduces the first; refusal reduces only the second. A discipline that refuses everything has full decision coverage and none of the second, which is the problem statement's trivially safe and practically worthless stated precisely, and the earlier version of this section placed it at zero coverage, which was wrong under either definition.
So the evaluation reports predeclared operating points rather than a curve. Four quantities carry it: decision coverage; warranted-practice retention, the proportion of independently warranted acts the discipline admitted, which is where over-conservatism becomes visible; unsafe-admission risk, the proportion of admissions that were not warranted, which is the closest equivalent to selective risk; and the undetermined rate, broken down by the axis that could not be evaluated, so that a gap in plant instrumentation is not read as a failure of discrimination.
Two disciplines govern their use. Unsafe-admission risk is undefined where nothing was admitted and is reported as undefined, never as zero, because a discipline that refuses everything would otherwise appear perfectly safe; it is warranted-practice retention that exposes it. And unsafe admissions and excessive refusals are not summed into one error rate, because their consequences are asymmetric and the weighting between them is a judgement this work is not entitled to make on a plant's behalf. Points may be compared across baselines, and a frontier reported where variants are non-dominated, but points are not joined into a line unless one genuinely ordered parameter is varied under a fixed reference task, such as a numeric freshness tolerance within a single act class.
What this changes is the kind of ground truth the evaluation needs, which an earlier statement of it overstated. That statement claimed the evaluation needs an outcome for each scenario, which is whether acting was in fact warranted. That sentence equates two things the work depends on keeping apart. An outcome is what happened after the act; a warrant is whether the basis available beforehand justified it. They come apart in both directions, and the corpus already relies on their coming apart: a claim may be true, and acted upon successfully, on a basis that did not justify the act, which is the structure the epistemology of testimony supplies and which this work cites for the derivation rule. An evaluation that read success as warrant would count a lucky unwarranted act as a correct admission, and would be measuring consequence while claiming to measure discrimination. The distinction matters because the alternative, ranking cases by a confidence value the model emits, is ruled out by the review's finding that estimating predictive uncertainty is not sufficient for safe decision-making (Gawlikowski et al., 2021): a curve read over a signal of the wrong category would measure the discipline against precisely what it exists to replace. So the evaluation needs two variables where it assumed one. A reference warrant states whether the basis available before the act justified it, and it is a normative judgement that cannot be observed. An outcome states what followed, and it can be. The corrected commitment is not that the designer avoids judgement but that the designer does not supply it: a reference warrant is grounded, wherever a case allows, in a requirement authored outside this work, in a standard, an operating procedure, a permit condition or a recorded engineering rule, with its source, the information available at the time, and any disagreement recorded alongside it. Calling the result a reference warrant rather than ground truth is deliberate, because it says what its epistemic status is.
Three evaluations, not one
Separating warrant from outcome separates the evaluation into three, and running them together would let a result in one stand as a result in another. Each answers a different question and admits different evidence.
Conformance asks whether the implementation does what the discipline says. It is answered by tests over stated properties, truth tables for the verdict, metamorphic cases and independently computed comparisons. It needs no experts and no outcomes, and it establishes only that the artefact demonstrated is the artefact designed.
Discrimination asks whether the discipline separates cases that operational practice already treats differently. It is answered against reference warrants, grounded outside this work where the case allows. This is where coverage and risk belong, as predeclared operating points rather than a curve, and it is the question the work is falsifiable on.
Consequence asks what happens when the discipline is used: unsafe acts prevented, legitimate acts refused, claims left undetermined, delay introduced, configuration effort, and the pressure to bypass a gate that stops useful work. This is where outcomes belong and where they are evidence rather than a proxy for something else.
One circularity survives the split and is guarded separately. If the same demand both generates a scenario's expected verdict and is supplied to the artefact under test, the exercise shows only that the implementation applies its own rules. So a reference warrant is stated first in domain terms, in the language its source uses, and translated into a demand afterwards, by a step that is recorded. That keeps two questions apart: whether the demand language can represent the domain requirement at all, and whether the comparison then reproduces the reference judgement. A failure in the first is a finding about the model, a failure in the second is a finding about the discipline, and a design that conflated them could show perfect agreement while representing nothing.
The residual, stated rather than hidden
One residual is real and is named here rather than smoothed over, because a section that claimed the objection fully dissolved would be committing the chapter's own cardinal sin. The scenario set is still the designer's. The individual verdicts are grounded in reference warrants rather than in the designer's opinion, but the distribution of cases is authored, and a set quietly weighted toward easy cases would flatter the operating points without any single verdict being false.
Three things bear on that residual, and none of them is a claim to have removed it. The scenarios are fixed before the rules harden, so the distribution at least cannot be fitted to the discipline after the fact. The composition criteria of the earlier section test whether what was built is the missing thing at all, independently of how it scores on any distribution, so a flattering result over a well-chosen set cannot rescue a discipline that is really a neighbour in disguise. And the distribution itself is declared and argued in the open, because a scenario set whose provenance is stated is a scenario set that can be attacked, and being attackable on that point is the honest position rather than a solved one.
What is still owed, and what nearly happened instead
This is a route rather than a settled answer, and saying so is part of keeping it honest. It needs the selective-prediction literature re-read on its evaluation protocol specifically, rather than on its contribution, which is a different reading from the one already done. And it needs a defensible account of what error means when the quantity of interest is not a misclassification but an action taken on a value whose basis did not warrant it. Both are tractable, and neither is the thing that was feared, which was that the discrimination question had no answer that did not route through the designer's judgement.
It is worth recording what nearly happened here, because this chapter exists to show its working and this is the sharpest instance of the working going one way rather than another. The move available a step before this one was to declare the designer's judgement an unavoidable limitation, observe that the discipline can still remove human vouching from every handover except its own bootstrap, and present that as the contribution rather than as a gap. It is a tidy move and it reads as candour, which is exactly why it is dangerous. It was wrong, and the answer that showed it was wrong was sitting in the project's own bibliography, read months earlier and filed under a different question. Declaring a problem unsolvable is a claim like any other, and it turns out to need a basis too. The discipline the whole work is about applied, in the end, to a claim the work was tempted to make about itself.
References
Fletcher, S. C. (2024). Reproducibility of Scientific Results. Stanford Encyclopedia of Philosophy. plato.stanford.edu/entries/scientific-reproducibility
Mandl, M. M., Becker-Pennrich, A. S., Hinske, L. C., Hoffmann, S. and Boulesteix, A.-L. (2024). Addressing Researcher Degrees of Freedom Through minP Adjustment. arXiv:2401.11537. arxiv.org/abs/2401.11537
Geifman, Y. and El-Yaniv, R. (2019). SelectiveNet: A Deep Neural Network with an Integrated Reject Option. ICML 2019. arXiv:1901.09192. arxiv.org/abs/1901.09192
Mozannar, H. and Sontag, D. (2020). Consistent Estimators for Learning to Defer to an Expert. ICML 2020. arXiv:2006.01862. arxiv.org/abs/2006.01862
Gawlikowski, J., Njieutcheu Tassi, C. R., Ali, M., Lee, J., Humt, M., Feng, J., Kruspe, A., Triebel, R., Jung, P., Roscher, R., Shahzad, M., Yang, W., Bamler, R. and Zhu, X. X. (2021). A Survey of Uncertainty in Deep Neural Networks. Artificial Intelligence Review. arXiv:2107.03342. arxiv.org/abs/2107.03342