Skip to content
Paula Livingstone writing · projects · tools

Attestable Methodology: First Draft (superseded)

Criteria the Design Did Not Author

The circularity problem needed criteria the designer did not write, and the review left them behind without meaning to. Every field it surveyed stated, in its own words and its own specification, the point at which it stops: the provenance standard locates assessment outside itself, the attestation framework declines to validate what it carries, the uncertainty survey calls its own estimates insufficient for safe decision-making. Quoted rather than paraphrased, these become a test authored by people who never heard of this work, and the test is falsifiable: if the basis model reduces to a provenance record that hands the judgement on, the provenance standard refutes it. The section closes on its own boundary, since these criteria test whether the missing composition was built and say nothing about how well it discriminates, which is a different question deferred to the evaluation.

Superseded. This is an earlier draft of the methodology chapter, retained as a record of what the work believed at the time. It is not current: the composition it describes includes constrained elevation, which the design chapter withdrew after implementation established that no axis of a basis is elevatable. The current chapter is the one published under the methodology category.

The previous section named the circularity problem and committed the chapter to answering it in two parts, one of which was that as much of the evaluation as possible be inherited from outside the design. This section makes good on that commitment. It does so with an asset the review produced without setting out to, and which is easy to walk past: as it read each field in that field's own primary source, it recorded, in each field's own words, the point at which that field stops.

Those stopping points are not a catalogue of failures. Each is a field stating the boundary of what it was built to do, usually as a design virtue rather than an omission, and stating it in a specification written by people who had never heard of this work and had no stake in its result. That is exactly the property the circularity problem requires. A criterion the designer authored is the designer's judgement wearing a different hat; a criterion quoted from a standard is not. So the discipline of this section is strict and worth naming before the criteria themselves. Each stopping point below is either quoted verbatim from the field's own specification or reported faithfully from it with a citation, and where it is reported rather than quoted it is not dressed in quotation marks, because a paraphrase in quotation marks claims a fidelity it does not have and is worse than an honest report. What must not happen is the boundary becoming the designer's account of the field a second time over, which would forfeit the one thing that makes the criterion independent: that the field, not the designer, drew the line.

The stopping points, in the fields' own words

Provenance locates the assessment outside itself. The W3C standard's own overview says of the relevant construct that it "describes a potential use of the recorded information by external parties, not an automated function within PROV itself" (W3C, 2013). The record is complete and the judgement is elsewhere, by the standard's own account.

Attestation declines to validate what it carries. in-toto, by its specification, does not validate whether the claims it carries are factually correct; it binds and authenticates the metadata and leaves the truth of the predicate alone (in-toto, n.d.), and SLSA that it "makes no claims about whether produced artifacts are functionally correct, true, or fit for purpose" (OpenSSF, 2023). The signature binds custody and says nothing about basis, by design and by the specification's own statement.

Access control has a slot for the judgement and does not fill it. The attribute-based access control guidance allows a confidence or assurance measure to be supplied to it as an attribute and factored into a decision, but it does not itself produce such a measure (Hu et al., 2014). The engine evaluates whatever assurance it is handed; where that assurance comes from is outside its scope, on its own terms.

Uncertainty quantification calls its own output insufficient for the decision. The field's large survey concludes that "estimating the predictive uncertainty is not sufficient for safe decision-making" (Gawlikowski et al., 2021). The number the field produces is, by the field's own capstone judgement, not the thing a safe decision needs.

Machine-learning assurance requires the sufficiency criterion as an input rather than producing it. The AMLAS methodology "requires as input the system safety requirements," which are "expected to be generated by domain experts or derived from the relevant regulatory requirements" (Hawkins et al., 2021). The argument it builds is only as good as a standard supplied to it from elsewhere, which is the standard admissibility would have to be.

Operational-technology security names the failure and treats it as a hazard to be prevented rather than a claim to be adjudicated. The NIST guidance lists, among the ways such a system can be harmed, the sending of inaccurate information to system operators, whether to disguise unauthorized changes or to cause operators to initiate inappropriate actions (Stouffer et al., 2023). The field sees a value that misleads the recipient into an unsafe act, and its response is to keep the value out, not to represent whether the value's basis warranted the act.

Turning quotations into a test

Quoted this way, the stopping points do more than corroborate the review. They compose into a criterion the artefact can be held against, and the criterion is falsifiable rather than rhetorical. Each field says, in effect, that whatever the artefact is, it must not claim to do what this field explicitly states cannot be done by its means. Three of these are concrete enough to fail the work outright, and stating them as failures is what makes them a test rather than a restatement.

If the basis model reduces to a provenance record that faithfully documents origin and hands the judgement onward, the provenance standard refutes it, because that is precisely the use by external parties PROV locates outside itself (W3C, 2013a), and the work would have rebuilt PROV and called it something new. If admissibility reduces to a confidence threshold, a value accepted when its number clears a bar, the uncertainty survey refutes it, because the field that produces that number states it is not sufficient for safe decision-making, and the work would have adopted as its answer the very thing the source field disowned. If the sufficiency criteria must be supplied to the mechanism from outside, then the mechanism is an assurance case with extra steps, and AMLAS refutes it (Hawkins et al., 2021), because that is what an assurance case already is: an argument over safety requirements handed to it as input.

These are the failures the review left the design exposed to, and they are worth stating plainly as such: the artefact fails if it collapses into any one of them. Put positively and once, the criterion is that the artefact must not reduce to any of the fields the review distinguished it from. The review's whole argument, field by field, becomes the standard the design is measured against, and the standard's authorship is entirely independent of the design because every clause of it is a quotation.

What these criteria do not test

This is where the section has to state its own boundary, because inherited criteria are easy to overclaim and the overclaim would quietly reintroduce the confusion the typing section removed. What these criteria test is the validity of the composition: whether the thing built is the missing composition the review identified, or a near-neighbour dressed in new terms. That is a real and necessary test, and it is the one circularity most threatened, because it is the judgement most tempting for the designer to make in the design's favour.

It is not a test of discrimination. Whether the discipline rejects False Determinism without rejecting so much legitimate practice as to be unusable is a different question, a property of the whole discipline rather than a matter of which field it does or does not resemble, and no quotation from a neighbouring field can answer it. A basis model can satisfy every criterion above, resembling none of the fields it was distinguished from, and still discriminate badly, refusing everything or accepting too much. The inherited criteria would pass it and the discipline would still be worthless. Discrimination is what the fourth research question asks, it is the question the work is falsifiable on, and it is answered by a different instrument, deferred to the evaluation and set out there.

Naming that boundary is not a hedge. It is the same distinction the typing section drew when it took the discrimination question out of the instantiation's row, and it is drawn here for the same reason: to keep two separate tests from being read as one. The criteria of this section establish that the right thing was built. Whether the right thing works is the evaluation's burden, and it is not discharged here and does not pretend to be.

References

W3C (2013). PROV-Overview: An Overview of the PROV Family of Documents. W3C Working Group Note. w3.org/TR/prov-overview

in-toto (n.d.). in-toto Attestation Framework: Specification. github.com/in-toto/attestation

OpenSSF (2023). Supply-chain Levels for Software Artifacts (SLSA), v1.0 Specification; Threat Model. slsa.dev/spec

Hu, V. C., Ferraiolo, D., Kuhn, R., Schnitzer, A., Sandlin, K., Miller, R. and Scarfone, K. (2014). Guide to Attribute Based Access Control (ABAC) Definition and Considerations. NIST Special Publication 800-162. doi.org/10.6028/NIST.SP.800-162

Gawlikowski, J., Tassi, C. R. N., Ali, M., Lee, J., Humt, M., Feng, J., Kruspe, A., Triebel, R., Jung, P., Roscher, R., Shahzad, M., Yang, W., Bamler, R. and Zhu, X. X. (2021). A Survey of Uncertainty in Deep Neural Networks. arXiv preprint arXiv:2107.03342. arxiv.org/abs/2107.03342

Hawkins, R., Paterson, C., Picardi, C., Jia, Y., Calinescu, R. and Habli, I. (2021). Guidance on the Assurance of Machine Learning in Autonomous Systems (AMLAS). University of York. arxiv.org/abs/2102.01564

Stouffer, K., Pease, M., Tang, C., Zimmerman, T., Pillitteri, V., Lightman, S., Hahn, A., Saravia, S., Sherule, A. and Thompson, M. (2023). Guide to Operational Technology (OT) Security. NIST Special Publication 800-82 Revision 3. csrc.nist.gov/pubs/sp/800/82/r3/final