Superseded. This is an earlier draft of the methodology chapter, retained as a record of what the work believed at the time. It is not current: the composition it describes includes constrained elevation, which the design chapter withdrew after implementation established that no axis of a basis is elevatable. The current chapter is the one published under the methodology category.
The inherited criteria of the previous section test whether the thing built is the missing composition rather than a neighbour in new terms. They do not say whether it works, and a discipline can satisfy every one of them and still be useless. This section states, before the artefact exists, what working and failing would mean, and it does so in advance for the same reason the scenarios are fixed in advance: a criterion written after the result is a criterion fitted to it.
Declaring success is the easy half and the less important one. The discipline succeeds to the extent that it refuses claims whose basis does not warrant the intended act while admitting those whose basis does, at the machine handover where no person is standing to make that call. The weight of this section is on failure, because failure is where a design-science evaluation is usually too generous with itself, and because in this case one of the two ways to fail is the one most easily disguised as caution.
Two faces, and why naming only one is a test that cannot be failed
The problem statement already fixed the shape of this, and did so in a single line: a discipline that rejects everything is trivially safe and practically worthless. The same asymmetry is visible in the field this work borrows its evaluation instrument from, where a classifier with a reject option is measured by what it achieves at a given coverage rather than by how much it declines, precisely because declining everything is available to any system and worth nothing (Geifman and El-Yaniv, 2019). That sentence is the whole of this section compressed, and unpacking it is the work. It says that failure is not one thing. There is the failure everyone expects a safety mechanism to have, and there is a second failure that a safety mechanism is uniquely tempted to hide, and a chapter that declares the first and treats the second as a caveat has quietly defined a test it cannot fail.
The first face is the false positive: the discipline admits a claim it should have refused, a case of False Determinism passes the boundary, and a value whose basis did not warrant the act is acted upon. This is the failure the whole work exists to prevent, and it is the one no reader will forget to check for.
The second face is the false negative: the discipline refuses a claim it should have admitted, and does so widely enough that legitimate practice cannot proceed through it. A basis that genuinely warranted the act is turned away, and if that happens often enough the discipline is not deployed, because a boundary that stops the plant running is removed by the people running the plant. This failure is quieter, and it is the one a conservative mechanism slides into while looking responsible.
The reason both must be declared, with equal weight, is arithmetic before it is ethical. A discipline that refuses everything has a false-positive rate of zero. It never admits a bad claim because it never admits any claim. If the only declared criterion of failure is admitting False Determinism, then refusing everything is a perfect score, and the evaluation has been arranged so that the laziest possible discipline wins it. The problem statement saw this exactly and said so: the work stands or falls on the discrimination it achieves, not merely on its conservatism. A test that conservatism alone can pass is not a test of the thing the work claims to do.
Over-conservatism is a failure mode, not a safe harbour
This has a consequence worth stating on its own, because it runs against the grain of how safety work usually reasons. In most safety settings, refusing when in doubt is the correct default, and a mechanism that errs toward refusal errs on the safe side. That is not true here, and the difference is the point. The discipline's purpose is not to maximise refusal; it is to make the right refusal and the right admission, because both a missed refusal and a needless one are ways of getting the boundary wrong. A discipline tuned to refuse whenever it is unsure would satisfy the safety intuition and defeat the contribution, because it would have rebuilt the thing every surveyed field already offers, a boundary a human removes because it stops useful work, rather than the thing the work claims, a boundary that discriminates. That over-refusal is defeated rather than tolerated is not a supposition. The human-factors literature names the pattern directly as disuse, the rejection or disabling of automation whose alarms or refusals the operator has learned to discount, and sets it alongside misuse as one of the ways reliance goes wrong (Parasuraman and Riley, 1997). A gate that refuses too often is not left in place refusing; it is worked around, and the discipline then governs nothing.
So over-conservatism is declared here as a failure, not a fallback. A version of the discipline that is safe because it refuses too much has failed, and it has failed in a way this chapter commits in advance to counting as failure rather than excusing as prudence. This is the declaration that gives the false-negative face its equal weight: it is not enough to say the discipline should not reject too much; the chapter must be willing to call a discipline that rejects too much a failed one, and it does.
What this section declares, and what it does not
What is fixed here is the definition of the two failures and the commitment to weigh them equally. That is the axis-setting: the discipline is judged on both whether it admits what it should refuse and whether it refuses what it should admit, and neither face is a footnote to the other. Fixing this before the artefact exists is what stops the criteria being fitted to the result, and stating both faces with equal weight is what stops the criteria being a test conservatism can pass by refusing to play.
What is not done here is the measuring. Declaring that over-conservatism counts as failure is not the same as showing how much is too much, and the two faces stated above are a pair of axes rather than a reading on them. The reading is the discrimination question, whether the discipline separates the claims it should refuse from those it should admit across the range between refusing everything and admitting everything, and that question is answered by a specific instrument, the coverage and risk relationship borrowed from the fields that already evaluate systems whose job is to decline. That instrument, and the scenarios it runs against, belong to the evaluation and are set out there.
The boundary is kept clean here for the same reason it was kept clean when discrimination was taken out of the typing table and again when the inherited criteria were held to testing composition rather than performance. This section says what would count as failing. It does not say how the failure is detected, because a chapter that mixed the declaration of the criteria with the method of measuring them would blur exactly the line that keeps the designer from setting the test and passing it. The criteria are declared now, in advance and in the open. Whether the artefact meets them is measured later, against them, by an instrument that was not built to flatter them.
References
Parasuraman, R. and Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors, 39(2), pp. 230-253. doi.org/10.1518/001872097778543886
Geifman, Y. and El-Yaniv, R. (2019). SelectiveNet: A Deep Neural Network with an Integrated Reject Option. ICML 2019. arXiv:1901.09192. arxiv.org/abs/1901.09192