Skip to content
Paula Livingstone writing · projects · tools

Attestable Methodology: First Draft (superseded)

Warranting Without Vouching: The Problem This Chapter Solves

The review closed by asking whether the missing composition can be built. An absent thing is simply a thing to build; what the review makes hard is the warranting. Its finding was that every surveyed field ends by handing the judgement to a human who can ask, so a discipline shown to work because its designer judged it works has placed a human at the very boundary the thesis says must be mechanised. The chapter's task: establish that a discipline for refusing unwarranted claims is itself warranted, without standing where the review said no-one is standing. Design science follows from the review rather than being chosen ahead of it; Peffers supplies the entry point and the artefact vocabulary.

Superseded. This is an earlier draft of the methodology chapter, retained as a record of what the work believed at the time. It is not current: the composition it describes includes constrained elevation, which the design chapter withdrew after implementation established that no axis of a basis is elevatable. The current chapter is the one published under the methodology category.

The preceding chapter ended on a question rather than a conclusion. Having walked ten literatures and established that none supplies, as a single working capability, a portable and per-claim representation of evidential basis that supports action-relative admissibility, derivation-aware inheritance, constrained elevation, and mechanical enforcement across an ownership boundary, it closed by asking whether such a thing can be built. This chapter is the answer to a narrower question that has to be settled first: how would anyone know?

That is not the question methodology chapters usually take, and the difference is worth stating at the outset. An artefact that does not exist is simply a thing to build, and nothing in the review makes the building difficult. What the review makes difficult is the warranting, and it does so in a specific way that the rest of this chapter is arranged around.

The trap the review sets for this chapter

The review's most consequential finding was not that the surveyed fields stop short of admissibility. It was where they stop. Every one of them, having built what it was built to build, hands the remaining judgement to a human being who can, in principle, ask what a value rests on. The provenance record informs a user who decides, its own standard describing provenance as information used to form assessments of quality, reliability or trustworthiness rather than as the assessment itself (W3C, 2013a). The datasheet advises a person choosing whether to deploy a model (Gebru et al., 2018). The safety case is a document that persuades a reviewer, and the guidance requires the system safety requirements as an input generated by domain experts (Hawkins et al., 2021). The governance mandate is owed to a deployer so that they may interpret an output and use it appropriately (European Union, 2024), and to those affected so that they may understand and challenge it (OECD, 2019). The human-factors literature, after forty years, concludes that the operator must calibrate their reliance, which requires understanding the automation's purpose, process and performance (Lee and See, 2004). The epistemology of testimony asks whether the hearer is entitled to believe (Leonard, 2021). Ten fields, and each ends with a person at the boundary.

The contribution's whole claim is that this is not good enough where no person is standing. That claim has an immediate and awkward consequence for the chapter now beginning. If the discipline is shown to work because its designer judged that it works, then the work has done precisely what it criticises: it has placed a human at the boundary to vouch for something whose basis was never independently carried. The objection does not need to be constructed by a hostile reader; it follows from the review's finding by a step any reader can take. The review establishes where the fields stop. The contribution claims that stopping there is insufficient. A methodology that then rested on the designer's judgement would stop in the same place, and the objection is the more damaging for being the thesis's own argument turned around. The requirement below is therefore inferred from the review rather than stated by it, and the inference is set out here so that it can be refused rather than assumed.

So the chapter's task can be stated exactly:

Establish that a discipline for refusing unwarranted claims is itself warranted, without relying on the designer's judgement as the evidence for its validity.

Every commitment in this chapter follows from that sentence, and any part of the chapter that does not serve it is ballast. The two that matter most are that the criteria for success are fixed before the artefact exists to be flattered by them, and that as many of those criteria as possible are authored by somebody else.

Why design science

The paradigm follows from the review's finding rather than being chosen in advance of it. Where a phenomenon exists, it can be observed, surveyed, measured, or experimented upon, and the paradigms that do those things are available. The review established that the composition at issue does not exist. There is nothing to observe. The only way to find out whether it can be built is to build it and see what the building teaches, which is design science's defining condition: where the natural and social sciences try to understand reality, design science attempts to create things that serve human purposes (Peffers et al., 2007).

The methodology this work follows is Peffers' Design Science Research Methodology (Peffers et al., 2007), which is the recognised process framework for such work and supplies a structure a reader can hold the research to. Two of its features do load-bearing work here and are named for that reason rather than for completeness.

The first is its treatment of entry points. The six activities of the process, problem identification and motivation, definition of objectives, design and development, demonstration, evaluation, and communication, are described as a nominal sequence rather than a mandatory one, and the paper is explicit that research may begin at almost any step and move outward, identifying four entry points by which it commonly does (Peffers et al., 2007). This work enters at the first, problem-centred, which matters because it means the literature review was not a preface to the research. It was the first activity of it: the problem identified, and its importance shown, from the surveyed fields' own statements of where they stop.

The second is the paradigm's definition of what an artefact is, which supplies this chapter's vocabulary rather than requiring it to invent one. Design science artefacts are constructs, models, methods, or instantiations (Peffers et al., 2007), and the next section reads the review's composition into those four categories. That the composition maps onto them without strain is worth a moment's notice: it is some evidence that what the review found is a design problem rather than a philosophical one, and therefore that a design-science answer is the right kind of answer.

Beyond those two points the paradigm is named, positioned, and left. Design science research is not a technique that helps anyone build anything, and a chapter that spends its length justifying its choice of paradigm has spent its length on the part of the work nobody contests. What is contestable is what gets built, how it will be judged, and by whose criteria, and the rest of this chapter is about those.

The shape of what follows

The chapter proceeds from what is being built to how it will be shown to have failed. It first types the artefact, reading the review's composition into constructs, a model, methods, and an instantiation, and states what each part must represent, because a description too vague to implement is also too vague to argue with. It then names the central methodological threat, which is the circularity of designing a discipline and its tests together, before setting out the two commitments that answer it: scenarios fixed before the rules harden, and criteria inherited from the fields the review surveyed. It states in advance what would count as success and, more importantly, what would count as failure, of which there are two kinds and only one is usually admitted. It schedules a debt the review incurred in public and would otherwise be free to forget. It closes on the limitations that remain when all of that is done.

References

W3C (2013a). PROV-DM: The PROV Data Model. W3C Recommendation. w3.org/TR/prov-dm

Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daume III, H. and Crawford, K. (2018). Datasheets for Datasets. arXiv:1803.09010. arxiv.org/abs/1803.09010

Hawkins, R., Paterson, C., Picardi, C., Jia, Y., Calinescu, R. and Habli, I. (2021). Guidance on the Assurance of Machine Learning in Autonomous Systems (AMLAS). University of York. arXiv:2102.01564. arxiv.org/abs/2102.01564

European Union (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 13. Official Journal of the European Union. eur-lex.europa.eu

OECD (2019). Recommendation of the Council on Artificial Intelligence, Principle 1.3. OECD/LEGAL/0449. legalinstruments.oecd.org

Lee, J. D. and See, K. A. (2004). Trust in Automation: Designing for Appropriate Reliance. Human Factors, 46(1), pp. 50-80. doi.org/10.1518/hfes.46.1.50_30392

Leonard, N. (2021). Epistemological Problems of Testimony. Stanford Encyclopedia of Philosophy. plato.stanford.edu/entries/testimony-episprob

Peffers, K., Tuunanen, T., Rothenberger, M. A. and Chatterjee, S. (2007). A Design Science Research Methodology for Information Systems Research. Journal of Management Information Systems, 24(3), pp. 45-77. doi.org/10.2753/MIS0742-1222240302