Superseded, retained as a record of method. This is the scenario protocol as published before any case existed, and it is not part of the evaluation chapter. The evaluation actually carried out departs from it in several respects, each a correction rather than a drift: warrant and outcome are separated, where this document still speaks of grounding verdicts in what followed the act; the artefact was frozen before scenarios existed, rather than scenarios being fixed before the rules hardened; the domain and translation stages were sealed separately; the source register was widened and every locator verified; and reference warrants requiring an inferential bridge this work supplied are marked as stratum 2 and reported separately, where this document holds that the designer should not supply the judgement at all. That last is the sharpest departure and the most honest: the bridges exist, so they are declared rather than denied. The protocol as executed is described in the evaluation chapter.
The methodology chapter made two commitments this stage has to keep. The scenarios are fixed before the admissibility rules harden, so the rules cannot be fitted to the cases. And the discrimination question is answered by outcome rather than by the designer's verdict, so the ground truth is observed rather than authored. The instrument that makes this workable is borrowed from the field that already evaluates systems permitted to decline, where performance is read as a trade of coverage against the error rate on what was accepted (Geifman and El-Yaniv, 2019). This document sets the protocol that keeps both, and it is published deliberately empty of cases.
The reason for publishing the protocol before any case exists is specific to what this stage is. A scenario set is the easiest place in the entire dissertation to cheat without noticing. Cases chosen a little too clean, a distribution weighted quietly toward what the discipline does well, expected verdicts that record the author's hopes rather than a prediction that could be wrong, and the whole thing still reads as rigorous. None of those are caught by inspecting a finished set, because a finished set that has been tilted looks exactly like one that has not. They are caught by fixing the method for building the set, in the open, before the building starts, and letting the method be attacked while changing it is still free. Once cases are drafted, the effort sunk into them makes the method harder to tear up, which is precisely when a reader most needs to have already agreed it.
So this spine names the protocol and lists no cases. That restraint is not tidiness; it is the point. A scenario protocol that quietly becomes scenarios has closed its own attack window, and the value of publishing now is entirely in the window staying open.
The questions the protocol must answer
Five questions decide whether a scenario set is honest, and the protocol has to answer all five before a case is written. Each is stated here as a commitment to a method, not as an answer about any particular case.
Where the cases come from
The provenance of the cases is itself a claim the set makes, and it has to be sourced rather than invented. The commitment is that the scenarios are derived from operational-technology situations described in the surveyed and cited literature and in public standards, rather than composed freehand to suit the discipline. The sources that qualify are those the review already read: the authoritative operational-technology guidance, which catalogues the ways such a system can be harmed and names the sending of inaccurate information to operators among them (Stouffer et al., 2023); the public incident record, of which the intrusion into a petrochemical safety-instrumented system is the documented case this work has already bounded to illustration (INL/DOE CyOTE, 2022); and the survey literature mapping threats and defences across the reference model's levels (Bhamare et al., 2020). A situation drawn from those can be checked by a reader against its source. A case whose situation can be traced to an external description is a case whose realism can be argued about by someone other than the author. The protocol fixes that cases must carry that trace; it does not here select which situations, because selecting them is drafting.
How an expected verdict earns authority independent of the author
This is the load-bearing question, because it is where the designer's judgement most wants to re-enter. An earlier version of this protocol tied a scenario's expected verdict to a determinate outcome, on the ground that an outcome is observed where a judgement is authored. The methodology has since withdrawn that, because an outcome is what happened after the act and a warrant is whether the basis justified it beforehand, and a lucky unwarranted act would otherwise be counted as a correct admission. What a case carries is therefore two things. A reference warrant states whether the basis available before the act justified it, and it is a normative judgement. An outcome states what followed, and belongs to the separate question of consequence rather than to discrimination.
The reference warrant is grounded outside this work wherever the case allows: in a standard, an operating procedure, a permit condition or a recorded engineering rule. Each case records the source of its criterion, the information available at the time of the act, who translated the source into the case, and any disagreement about it. Expert judgement is admissible here, which an earlier version denied. What is not admissible is the designer supplying the judgement, and those are different things: an accountable person authoring a criterion in advance is the same ex ante work the discipline already requires of whoever authors an act's demand. Where a reference warrant cannot be grounded in a source outside this work and the case is retained anyway, the case records that its warrant rests on judgement made for this evaluation, so that a reader can weigh those cases separately rather than discovering the difference in aggregate
How the distribution is kept honest
Individual verdicts can be outcomes and the set as a whole can still lie, if the mix of cases is weighted toward the discipline's strengths. The protocol keeps the distribution honest by three means, none of which is a claim to have removed the author from it. The mix is fixed before the rules harden, so it cannot be adjusted to the discipline after the fact. The mix is declared and characterised in the open, so that a reader can see its shape and attack it, a set whose composition is stated being one that can be challenged on composition. And the set is required to carry cases the discipline is expected to get wrong, so that it cannot consist only of cases built to pass. The residual that the distribution is still the author's is real, was named in the methodology, and is not claimed here to be dissolved.
How many cases, and of what kinds
The protocol fixes that the set spans both arms of the discrimination question in deliberate proportion rather than by whatever is easy to write. Both the cases that should be refused and the cases that should be admitted must be present in numbers that let coverage and risk be read against each other, since a set clustered at one end reports nothing about the other. What is produced is a set of operating points rather than a curve: a curve would require one genuinely ordered parameter varied under a fixed reference task, and this discipline has no such parameter. Joining fixed points with a line would assert a trade-off that was never traced. The kinds are fixed by what the discipline claims to distinguish, among them differences in the means by which a value was established, in the freshness of a basis against how fast its condition changes, in the scope a value was established within against the scope the act requires, and in the consequence of the intended act. The protocol commits to covering the kinds the discipline says it can tell apart; it does not enumerate the individual cases that populate them.
What an expected failure concretely means
The methodology declared failure in two faces and committed to including cases the discipline is expected to get wrong. The protocol makes that concrete as a category of the set rather than a rhetorical gesture. An expected failure is a case whose reference warrant is established and against which the discipline, as the author honestly predicts before the rules harden, will return the wrong verdict. Committing such cases in advance, with their expected wrong answer recorded, is what makes the evaluation falsifiable rather than a demonstration, because it creates results the author has predicted and can be shown to have gotten wrong. A set with no expected failures in it is a set built to pass, and the protocol treats the presence of pre-declared expected failures as a requirement, not an option.
What may not count as discrimination evidence
By the time this protocol is executed the artefact has a large body of tests behind it: an executable specification, its regressions, property tests over generated input, and a set of mutation operators that must all be killed. It would be easy, and wrong, to present that as evidence that the discipline discriminates.
Those cases shaped the model. Several of them exist precisely because the model got something wrong and was corrected, and a case written to record a correction cannot then be offered as an independent test of it. The evaluation therefore separates three classes of evidence and does not let one stand in for another.
Conformance evidence is the specification, the regressions, the properties and the mutants. It answers whether the implementation preserves the discipline the design chapter fixed. It is reported as conformance and claims nothing about discrimination.
Operational demonstration is the worked integration: a plant historian feeding a model output, adjudicated for two different acts, then re-established by an independent measurement and reevaluated. It was developed alongside the artefact and is therefore exploratory. It evidences mechanisability and that the pieces connect, and it is labelled as a demonstration wherever it appears.
Discrimination evidence is the confirmatory scenario corpus alone, authored after the artefact is frozen, from situations external to the design. It is the only class that speaks to whether the discipline tells warranted practice from unwarranted, and it is the class this protocol governs.
The corpus is therefore built fresh rather than assembled from the defects already recorded in the specification. Those remain conformance regressions and stay where they are.
The freeze, and what happens when evaluation finds a defect
Pre-registration as described above fixes timing. It does not on its own stop the more ordinary failure, which is that evaluation quietly becomes another development suite: a scenario fails, the model is adjusted, the corpus is rerun, and the published result is the second run with no record that the first happened.
So the artefact is frozen before the corpus is executed, and the freeze is captured mechanically rather than asserted. It records the library commit and a hash over its sources, the specification and mutation-set hashes, the schema and vocabulary versions, the protocol and evaluator versions, the metric definitions, the baseline definitions, and the claims the evaluation is capable of embarrassing. Recording the metrics before any result exists is what stops a metric being chosen to suit an outcome, and recording the claims is what stops a claim being narrowed to fit one.
The freeze also states what may change afterwards without invalidating results, and what may not. Prose, test names and the mutation runner's routing may change. The library's adjudication behaviour, the vocabulary and schema versions, the metric and baseline definitions, and any scenario's reference warrant or translation once execution has begun, may not.
When evaluation exposes a defect in the artefact, which it is expected to, the sequence is fixed: the original result is preserved and published, the failure is recorded as a finding, the artefact is revised, and the corpus is rerun as a separately identified version against a new freeze. Both results stand in the record. A defect found by evaluation is the most valuable thing the evaluation can produce, and overwriting it with the repaired run would discard exactly the evidence that the method works.
Every scenario names its decisive comparison
One further requirement, learned from the artefact rather than reasoned in advance. While building the worked demonstration an advisory demand was written that excluded an empty set of mechanisms. It excluded nothing, so the demand admitted for no reason at all, and the demonstration passed while proving nothing about the distinction it claimed to draw. Nothing in the result revealed this: a green test and a vacuous one look identical.
So a scenario records not only its expected verdict but the particular comparison that must be decisive: which axis, and which value against which requirement. A scenario that returns the expected verdict for the wrong reason is a failed scenario, not a passed one, and the trace is checked against the declared comparison rather than only the verdict against the expected verdict.
This is the same discipline the specification already applies to its negative cases, where a case must fail on the obligation it names rather than on any obligation at all. It is stated here because a scenario set is where a vacuous test is hardest to notice and most damaging.
Pre-registration, in operational terms
The commitments above all depend on a fixed before the rules harden that has to mean something checkable rather than being a claim the author makes about the author's own timing. The protocol's mechanism is that the scenario set, its expected verdicts, and its declared distribution are recorded and fixed as a unit before the admissibility rules are settled, and that this fixing is what the later evaluation is held to. The point of pre-registration here is narrow and worth stating exactly: it fixes timing, so the rules cannot be quietly fitted to the cases. It does not by itself fix authorship, which is why it stands alongside outcome-based verdicts and a declared distribution rather than in place of them. Pre-registration is necessary, it is not sufficient, and the protocol relies on it for exactly the one thing it can do.
What this spine is, and what it is not
What is fixed here is the method: the provenance of cases, the grounding of reference warrants outside this work, the three guards on the distribution, the coverage of both arms and the discriminated kinds, the concrete meaning of an expected failure, and the pre-registration that fixes the timing of all of it. What is deliberately absent is any case. No situation is described, no verdict is assigned, no distribution is populated, because the moment this document begins listing cases it has stopped being a protocol that can be attacked cheaply and has become a draft whose cost defends it. The scenario set is built against this protocol, next, once the protocol has been left open to attack. The line between naming the method and writing the cases is the line this stage most has to hold, and this spine holds it on purpose.
References
Stouffer, K., Pease, M., Tang, C., Zimmerman, T., Pillitteri, V., Lightman, S., Hahn, A., Saravia, S., Sherule, A. and Thompson, M. (2023). Guide to Operational Technology (OT) Security. NIST SP 800-82r3. doi.org/10.6028/NIST.SP.800-82r3
Idaho National Laboratory / U.S. Department of Energy CyOTE (2022). Case Study: TRITON Malware Attack Against Petro Rabigh. INL/RPT-22-67981. cyote.inl.gov
Bhamare, D., Zolanvari, M., Erbad, A., Jain, R., Khan, K. and Meskin, N. (2020). Cybersecurity for industrial control systems: A survey. Computers & Security, 89. arXiv:2002.04124. arxiv.org/abs/2002.04124
Geifman, Y. and El-Yaniv, R. (2019). SelectiveNet: A Deep Neural Network with an Integrated Reject Option. ICML 2019. arXiv:1901.09192. arxiv.org/abs/1901.09192