Between building the artefact and evaluating it, the paradigm places a further activity: running the artefact on instances of the problem so that its behaviour can be observed. Evaluation then compares those observations against the objectives fixed in advance (Peffers et al., 2007). This section states how that run is conducted.
What is run
The library is run, not the discipline. The claim under test is that admissibility can be decided at a machine handover with no person present, and a discipline applied on paper by its designer would test the opposite. So the artefact is executed, and what it decides is the observation.
It is run on the pre-registered scenarios. Those are fixed by a protocol settled before any case exists, which requires cases traceable to described operational-technology situations, verdicts grounded in a reference warrant — whether the basis available before the act justified it — rather than in the designer's expectation, and the inclusion of cases the discipline is expected to get wrong. The run therefore consumes a case set it cannot revise.
What is recorded
Five things per case. The basis record as it crossed the boundary. The act's demand as the case states it. The verdict returned, with the failing axis where the verdict is a refusal. The reference warrant: whether the basis available before the act justified it, a normative judgement grounded outside this work where the case allows. And, separately, the outcome: what followed the act.
The first three are what the artefact did; the last two are what the case supplies. Warrant and outcome are recorded as two quantities rather than one because they come apart in both directions — a true claim may be acted on successfully on a basis that did not justify the act — and collapsing them would measure consequence while claiming to measure discrimination. Keeping them apart is what lets a reader recompute a verdict from the record and the demand without repeating the run, which is possible because the comparison is over stated fields and the combination is conjunctive.
What happens when a case cannot be run
Some cases will not be expressible: a demand referring to something the record does not carry, or a situation the model cannot represent at all. These are recorded as coverage findings against the artefact and reported with the results.
They are not dropped. Dropping them would restrict the run to cases the model already handles and would convert an unfavourable result into a smaller sample, which is the failure the scenario protocol exists to prevent.
References
Peffers, K., Tuunanen, T., Rothenberger, M. A. and Chatterjee, S. (2007). A Design Science Research Methodology for Information Systems Research. Journal of Management Information Systems, 24(3), pp. 45-77. doi.org/10.2753/MIS0742-1222240302