Skip to content
Paula Livingstone writing · projects · tools

Attestable Methodology: First Draft (superseded)

The Circularity Problem, Named Before Its Answer

A discipline evaluated against cases chosen by the person who wrote its rules will pass, and the pass will mean nothing. That is the central methodological threat, and this section states it in the open before any of the safeguards that answer it, because a chapter presenting its defences without first naming what they defend against is asking to be taken on trust. Two commitments follow: scenarios fixed before the rules harden, and criteria inherited from outside the design. The first turns out to be the paradigm's own requirement rather than a scruple this work volunteered, which means a reader can hold the work to it without taking the author's word for why it matters.

Superseded. This is an earlier draft of the methodology chapter, retained as a record of what the work believed at the time. It is not current: the composition it describes includes constrained elevation, which the design chapter withdrew after implementation established that no axis of a basis is elevatable. The current chapter is the one published under the methodology category.

The preceding section said what is to be built and to what standard of precision. The moment a standard is named, a question follows that the rest of this chapter exists to answer: who applies it, and what stops the answer being flattered by the person with the most to gain from a good result. That person is the designer, and in a project with a single researcher the designer is also the evaluator. The threat this creates is the chapter's central one, and it is named here, in the open, before any of the commitments that answer it.

The threat is circularity, and it has a plain form. Design the discipline and the scenarios it will be judged against at the same time, by the same hand, and the discipline will pass. The scenarios will have been chosen, consciously or not, as the ones it handles well. The pass will be real in the sense that the rules did produce the verdicts, and worthless in the sense that the verdicts were arranged in advance. A reader has no way to tell a discipline that discriminates from one whose author simply knew which cases to bring, because from the outside the two produce the same table of results.

There is a reason to state this before its answer rather than after, and it is not only a matter of good manners. A methodology that presents its safeguards first and the threat second invites the reader to take the safeguards on trust, because the reader has not yet been shown what they are for and cannot judge whether they are sufficient. Naming the threat first fixes the standard the safeguards must meet. It also forecloses the more comfortable move of quietly choosing safeguards that happen to be easy to satisfy and presenting them as though the threat had been fully met. The threat stated plainly is harder to under-answer without it showing.

Why this threat is the thesis again

Circularity would be a serious problem for any design science project, and it is a more exact one here, because it is the same shape as the failure the work is about. The review's finding was that every surveyed field ends by handing the judgement to a human who can, in principle, ask what a value rests on. The contribution's claim is that this is not good enough where no person is standing. If the discipline is shown to work because its designer judged that it works, the evaluation has placed a human at the boundary to vouch for the very mechanism whose whole purpose is to remove the need for one. The work would be enacting, in its own evaluation, the failure it names in its subject.

This is why the circularity problem cannot be handled with the standard reassurances. It is not enough to promise care, or independence of mind, or to note that the researcher is aware of the risk. Awareness is exactly what the review found the surveyed fields already had: the machine-learning assurance guidance names automation bias and hands it on, the human-factors literature studies over-reliance for forty years and concludes the operator must calibrate. Awareness of the boundary was never the scarce thing. What the work owes is a way of establishing its result that does not route through the researcher's own judgement at the decisive point, and the two commitments below are the beginning of that.

The first commitment: scenarios before the rules harden

The first commitment is that the evaluation scenarios are fixed before the admissibility rules harden, and that they include cases the discipline is expected to get wrong. Fixing the scenarios first addresses the timing half of the circularity: rules cannot be quietly fitted to cases that were settled before the rules existed. Requiring that the set include expected failures addresses a subtler evasion, which is a scenario set consisting only of cases the discipline is confident about. A discipline demonstrated solely against cases it was built to pass has been demonstrated to do nothing. Committing in advance to particular cases the discipline should reject, and to particular cases it should accept, and being willing to be wrong about either, is what separates a falsifiable claim from a demonstration staged for the result.

This commitment turns out not to be a scruple this work invented, and that is worth establishing rather than asserting, because a commitment the author volunteered can be relaxed by the author, whereas one the paradigm requires cannot be relaxed without leaving the paradigm. Design science defines evaluation as comparing the objectives of a solution to actual observed results from use of the artefact in the demonstration (Peffers et al., 2007). The objectives are fixed at the second of the six activities, definition of objectives, and the artefact is designed and developed at the third. The ordering is the method's, not this project's: objectives precede design by construction. A design science project that settled its criteria after seeing how its artefact performed would not merely be exercising poor judgement. It would have stopped following the method it claims to follow, and the anti-circularity commitment would have been abandoned in the same motion.

The consequence is that a reader does not have to take the author's word for why fixing the scenarios first matters. The requirement is external to the author, published in the paradigm, and a reviewer can hold the work to it and check the ordering was kept. That is a stronger position than a promise, and it is available only because the discipline being followed already demanded the thing the circularity problem needs.

The second commitment: criteria authored elsewhere

Fixing the timing is necessary and not sufficient. Scenarios fixed before the rules still leave the authorship of the scenarios with the designer, and a set fixed early but chosen to be easy would satisfy the first commitment while defeating its point. The second commitment addresses authorship: as much of the evaluation as possible is inherited from sources outside the design, so that the standard the work is held to was written by people who never heard of it.

That the review has already supplied such sources is what makes this commitment concrete rather than aspirational, and setting them out is the work of the next section. What matters here is the shape of the answer to the circularity threat, which comes in two parts that meet its two halves. Timing is met by the paradigm's own ordering of objectives before design. Authorship is met, to the extent it can be, by taking the evaluation's criteria from the surveyed fields' own statements of where they stop, rather than from the designer's sense of what a good result would look like. Neither part is complete on its own, and one part of the threat, the authorship of the scenario distribution rather than the criteria, survives both and is dealt with directly when the evaluation is set out. It is named here so that the reader knows it has not been forgotten between the naming of the threat and the accounting for it.

What naming it first commits the chapter to

Stating the threat before the safeguards binds the rest of the chapter to a particular honesty. The success and failure criteria that follow must be ones the discipline could genuinely fail against, or the circularity has merely moved from the scenarios to the criteria. The inherited standard must be one the designer did not author, or the independence claimed for it is nominal. And the one residual that neither commitment closes must be reported as a residual rather than dissolved by a reassuring sentence. A chapter that named circularity as its central threat and then quietly declared victory over it would have demonstrated the threat rather than answered it. The point of naming it first is that the reader can watch whether it is actually met, and hold the chapter to the standard it set for itself before it knew whether it could reach it.

References

Peffers, K., Tuunanen, T., Rothenberger, M. A. and Chatterjee, S. (2007). A Design Science Research Methodology for Information Systems Research. Journal of Management Information Systems, 24(3), pp. 45-77. doi.org/10.2753/MIS0742-1222240302