Skip to content
Paula Livingstone writing · projects · tools

Attestable Evaluation

Discrimination Results: Five Cases, and What Five Cases Can Support

Execution ran against a sealed eligibility manifest naming only the exactly-classified cases, over the sealed domain and translation stages. The runner cannot reach a case outside the manifest, so that no approximate demand could quietly convert a representational failure into a discrimination figure.

Five cases were eligible under the frozen classification. Raw counts with denominators follow; with five cases no percentage is meaningful.

CaseExpectedActualOutcome
S-fieldprov-1: breaking containment, valve identity from a three-year-old spreadsheetrefuserefusecorrect, on the predicted axis
S-fieldprov-4: closing a valve, vessel attribution from the same unmaintained mappingrefuserefusecorrect, on the predicted axis
S-semantic-3: tripping a compressor on a verified figure behind an unverified conclusionrefuserefusecorrect, on the predicted axis
S-ancestry-1: adjusting a setpoint from an undocumented spreadsheet derivationrefuseadmitunsafe admission
S-ancestry-2: the same act with the generated origin recorded and visibleadmitrefuseexcessive refusal; translation defect

Three of five matched on both verdict and predicted decisive axis, as run under the sealed protocol. Among the four translations surviving exactness review, three of four. The second figure was not predeclared and is a sensitivity analysis rather than a result.

Each case also carried a prediction of which axis would decide it, made before execution. A case reaching the expected verdict on the wrong axis is recorded as a failure, because it has not tested what it claimed to test. All three matching cases decided on their predicted axis.

What these figures cannot support

None of the eligible cases is stratum 1. Every one rests on a warrant reached through an inferential bridge this work supplied between an external source and the requirement. Agreement therefore shows the artefact consistent with this work's own reasoning, not validated against established practice.

Three of the four validly exact cases are cases the artefact was specifically built to handle. The language represents the commitments this work formalised and rarely represents an externally stated requirement without a bridge of its own.

Three cases in a corpus of thirty-nine, all resting on authored bridges, cannot establish discrimination performance. The small eligible set arose principally from the representational result, although one eligibility classification was itself defective, and that defect is analysed in the next section.