Skip to content
Paula Livingstone writing · projects · tools

Attestable Evaluation

Threats to Validity

Stated as threats rather than as limitations, because each could change how the results should be read rather than merely bounding them.

Purposive allocation. The corpus was allocated deliberately across families and warrant arms, not sampled from any population. Counts describe this corpus. Nothing here estimates how often an operational requirement would prove unrepresentable in practice, and the representational failure rate should not be read as a rate.

Author-produced warrants. Reference warrants were fixed before execution and grounded in external sources, but eighteen of the thirty-nine cases required an inferential bridge this work supplied rather than a warrant stated outright by an external source, so eighteen carry this work's own reasoning in their ground truth. The concentration is sharper still at execution: every one of the four defensibly exact cases — the only ones any exact translation came from — is one of those stratum-2 cases, and no case whose warrant is stated outright by an external source was exactly representable, so the discrimination evidence rests entirely on cases carrying this work's own reasoning. Where a reviewer rejects a bridge, that case's result carries no weight. The eighteen bridges are recorded per case so that they can be rejected individually rather than in aggregate.

Author translation. The same person authored the requirements, the translations and the artefact. The staged seals prevent a requirement being adjusted after its demand is in view, and the atom decomposition prevents requirements being shaped around what the language can express. Neither prevents an author from misreading a source, and one translation defect survived every check until execution.

A small eligible set. Three matching cases support no general claim about discrimination. This is a consequence of the representational result rather than a design fault, but it means the second research question is answered only in the negative sense that it could not properly be asked.

No field users. No operational-technology engineer authored a demand, constructed a record, or read a verdict. Usability, comprehensibility and integration fitness are untested, and the construction-boundary failure above suggests they are where the artefact's practical viability will be decided.

Post-execution discovery. One classification defect was found after execution and is reported in both accounting views rather than corrected silently. Its discovery raises the possibility that other translations contain defects the checks did not catch. The translations are published in full so that this can be examined rather than taken on trust.

Single artefact version. The evaluation tests one frozen version. It establishes nothing about whether the identified losses are remediable, only that they were present.