Skip to content
Paula Livingstone writing · projects · tools

Attestable Process and Planning

The Spine: What the Methodology Must Do, and the Part That Is Not Solved

Published unfinished, on purpose. The review's last line is a commission: whether the missing composition can be built. That leaves one job, and the review's own finding sets the trap, since every field it surveyed ends by handing judgement to a human who can ask. A discipline validated by its designer's judgement puts a human back at the boundary the review said must be mechanised. The artefact's typing is inherited from the review rather than invented; failure is declared with two faces, since over-conservatism is a failure mode and not a safe harbour; the criteria are inherited from the ten fields' own stated limits, authored by people who never heard of this work. The discrimination problem, feared unanswerable without the designer vouching, has a route: the coverage and risk trade-off from the selective-prediction literature, which evaluates systems that decline without labelling any case as one that ought to have been refused.

This is a spine, not a chapter. It is published unfinished and deliberately so: the structure below is firm enough to attack, and one part of it is not solved at all. Setting it out in public is the cheapest way to find out where it breaks, and the alternative, refining it privately until it looks confident, mostly produces confidence rather than soundness.

What follows states what the methodology chapter must do, why the literature review and not a methods textbook decides that, the shape the chapter will take, and the one problem currently without an answer. Where the reasoning is still open it is marked open rather than smoothed over.

The commission

The literature review chapter ends on a sentence that is a commission rather than a summary: whether the composition it found missing can be built is the question the rest of this work asks. The review spent its length establishing, field by field, that a portable, per-claim representation of evidential basis, supporting action-relative admissibility, derivation-aware inheritance, constrained elevation, and mechanical enforcement across an ownership boundary, does not exist as a single working capability in any of the ten literatures that would be expected to hold it.

That leaves the methodology chapter with exactly one job, and it is not the job methodology chapters usually take. The job is not to name a paradigm and cite it. An absent thing is simply a thing to build, and nothing about the review makes the building difficult. What the review makes difficult is the warranting. The question it forces is this:

How do you establish that what you built is the thing the review found missing, and that it works, when your own judgement is the least admissible evidence available?

The trap, which is the same shape as the thesis

The review's most consequential finding is not that the fields stop. It is where they stop. Every one of the ten ends by handing the judgement to a human being who can, in principle, ask. The provenance record informs a user who decides. The datasheet advises a person choosing whether to deploy. The safety case persuades a reviewer. The governance mandate is owed to a deployer so that they may interpret and contest. The human-factors literature stops at the operator who must calibrate their reliance. The epistemology of testimony asks whether the hearer is justified. Each of those is a reasonable place to stop while a person stands at the boundary and can ask what a value rests on.

The consequence for this chapter is uncomfortable and unavoidable. If the discipline is validated by its designer's judgement, then the methodology has put a human back at precisely the boundary the review argued must be mechanised. The work would be demonstrating the failure it names, inside the chapter that claims to establish it. A critical reader does not need to construct that objection; the review hands it to them.

So the chapter's actual thesis is this: establish that a discipline for refusing unwarranted claims is itself warranted, without standing where the review just said no-one is standing. Everything below follows from that, and any part of the chapter that does not serve it is ballast.

Why design science, briefly

The review licenses the paradigm rather than the other way around. The artefact does not exist, so there is nothing to observe, survey, or experiment upon; the thing must be made in order to find out whether it can be made. That is design science's defining condition, and Peffers' Design Science Research Methodology is the recognised spine to hang it on (Peffers et al., 2007).

Two of its features do real work here rather than decorative work, so they are worth naming precisely. The first is that its six activities are described as nominally sequential while allowing four different entry points, since research may in practice start at almost any step and move outward. This work enters at the first, problem-centred, and the literature review is that activity rather than a preface to it: the problem was identified and its importance shown by survey, in the fields' own words. The second is that the paradigm's own definition of an artefact supplies the vocabulary the next section uses, so the typing there is the paradigm doing its job rather than a scheme invented for this chapter.

Beyond that the paradigm is named, positioned, and left. It is not a technique that helps build anything, and a chapter that spends its weight justifying its choice of paradigm has spent its weight on the part nobody contests.

Typing the artefact, inherited rather than invented

The review already did this work, which is worth saying plainly because it is what makes this chapter a continuation rather than a fresh start. The composition the review found missing is stated three times in the same words and then walked against the fields, each holding one or two of its parts. That is not a summary. It is a requirements list derived by survey, and the methodology reads it off rather than proposing its own.

The categories it is read into are not invented here either. Design science takes its artefacts to be "constructs, models, methods, or instantiations" (Peffers et al., 2007), and the composition maps onto those four without strain, which is some evidence that the review found a design problem rather than a philosophical one.

What the review found missing What kind of artefact that is
A portable, per-claim representation of evidential basis Constructs and a model
Derivation-aware basis preservation A method
Action-relative admissibility A method
Mechanical evaluation across an ownership boundary An instantiation

The discipline this imposes is precision. A basis model described as capturing "relevant provenance factors" is both weak prose and unimplementable; one that names what it represents can be argued with, and can be built. The chapter states what each part must represent and stops there. It does not sketch interfaces or type signatures, because at that point it stops being an account of how a claim gets made honestly and becomes engineering paperwork with a citation attached.

The circularity problem

The live methodological risk is designing the discipline and the test scenarios together. A discipline evaluated against cases chosen by the person who wrote the rules will pass, and the pass will mean nothing. This is stated here before its answer, because a methodology that presents its safeguards without first naming the threat they answer is asking to be taken on trust.

Two commitments follow. The first is that the evaluation scenarios are fixed before the admissibility rules harden, and include cases the discipline is expected to get wrong. Committing in advance to expected failures is what separates a falsifiable claim from a demonstration. The second is that as much of the evaluation as possible is inherited from sources outside the design, which is the next section.

The first commitment turns out not to be a scruple this work invented. The paradigm already requires it: evaluation in design science means "comparing the objectives of a solution to actual observed results from use of the artifact in the demonstration" (Peffers et al., 2007), and those objectives are fixed at the second activity, two steps before the artefact is designed at the third. A design science project that settled its criteria after seeing how its artefact performed would not merely be showing poor judgement; it would have stopped following the method it claims. That is worth stating plainly, because it means the anti-circularity commitment is inherited rather than volunteered, and a reader can hold the work to it without taking the author's word for its importance.

Criteria the design did not author

The review leaves behind an asset that is easy to miss. Every field it surveyed draws its own boundary, in its own words, in sources read in full and quoted: the provenance standard says assessment is a use by external parties rather than a function it performs; the attestation framework says it does not validate whether predicate claims are factually correct; the supply-chain framework makes no claims about fitness for purpose; the access-control standard has a slot for an assurance measure it does not produce; the uncertainty survey states that estimating predictive uncertainty is not sufficient for safe decision-making; the machine-learning assurance guidance requires its safety requirements as an input; the operational-technology guidance names inaccurate information reaching operators as a hazard.

Those are constraints authored by people who had never heard of this work, and they are testable. If the basis model turns out to be a provenance record that hands the judgement onward, the provenance standard refutes it. If admissibility reduces to a confidence threshold, the uncertainty survey refutes it. If the sufficiency criteria have to be supplied from outside the mechanism, then it is an assurance case with extra steps, and the assurance literature refutes it. Stated as one criterion: the artefact must not collapse into any of the ten fields the review distinguished it from. The review's argument becomes the evaluation's standard, and the standard's authorship is independent of the design.

Success and failure, declared in advance

The problem statement is already explicit that a discipline which rejects everything is trivially safe and practically worthless, and that the work stands or falls on the discrimination it achieves rather than on its conservatism. Failure therefore has two faces, and both are declared before the artefact exists to flatter: failing to reject False Determinism, and rejecting so much legitimate practice as to be unusable. Over-conservatism is a failure mode, not a safe harbour. A chapter that only defines the first kind of failure has defined a test it cannot fail.

The debt the review incurred

The review's synthesis conceded in public that several unsurveyed fields may already hold semantics this work would otherwise reinvent: evidence theory on corroboration, conflict, and dependence between items of evidence; decision theory on how sufficiency relates to consequence; trust management on delegated authority; contract-based design and runtime assurance on enforcement at a barrier; epistemic logic; proof-carrying code as a precedent for carrying justification alongside an artefact. It named that reading as owed before the basis model is fixed.

A debt named and not scheduled is a debt disowned, so the schedule is part of the method rather than an intention: draft the smallest basis model that could work, identify its load-bearing semantics, pressure-test only those against the fields above, and only then stabilise. The order matters in both directions. Reading six literatures speculatively before knowing which semantics are load-bearing is a way to spend months and learn little; freezing the model before the check is how you end up inventing a worse version of something that has existed since the 1960s.

The discrimination problem, and a route through it

The part the work is falsifiable on asks whether the discipline rejects False Determinism in realistic scenarios without rejecting so much legitimate practice as to be unusable. That is a claim about discrimination, and discrimination is only meaningful against a ground truth. Something has to say which scenarios ought to be rejected and which ought to pass, and the obvious candidates do not survive contact:

  • Pre-registration alone is insufficient. Fixing the scenarios and their expected verdicts before the rules harden solves the problem of timing: the rules cannot be quietly fitted to the cases afterwards. It does nothing about authorship. The designer's honest predictions are still the designer's.
  • Expert labelling relocates the problem. If domain experts supply the verdicts, a human still vouches, which is the thing the review identified as where every field stops.
  • Incident reports only cover half the question. They document what went wrong, and are already bounded in this work to illustration rather than evidence. The harder half is the legitimate practice that must not be rejected, and failures are written up where normality is not. Nobody publishes a report titled "the plant ran correctly on Tuesday."

The route through is that this problem is not new, and the field that solved it is one the review already surveyed. Selective prediction and learning to defer build systems whose job is to decline, and they face the identical question: how do you evaluate a mechanism whose output is sometimes a refusal? They answer it without labelling any individual case as one that ought to have been refused. The instrument is the coverage and risk trade-off: for each level of coverage, meaning the proportion of cases the system accepts, measure the error rate over what it accepted. The two degenerate positions need no judge at all, because they are visible on the axes. A discipline that refuses everything sits at zero coverage, which is the problem statement's "trivially safe and practically worthless" expressed as a coordinate. A discipline that accepts everything carries the base error rate unchanged and has added nothing. Discrimination is the shape of the curve between them, and it is read off rather than adjudicated.

What this changes is the kind of ground truth required. Not a verdict on each scenario, supplied by the designer, but an outcome for each scenario: whether acting on that value was in fact warranted. The discipline's ordering is then compared against the outcomes. The designer does not label what should be rejected; the designer builds cases whose outcomes are determinate, and the measurement follows.

One residual is real and is stated rather than hidden. The scenario set remains the designer's, so the distribution is authored even when the individual verdicts are not, and a set weighted toward easy cases would flatter the curve. Three things bear on that: the scenarios are fixed before the rules harden, so they cannot be fitted afterwards; the criteria of the previous section test whether what was built is the missing composition at all, independently of how it scores; and the distribution is declared and argued in the open, because a scenario set whose provenance is stated is a scenario set that can be attacked, which is the point of stating it.

This is a route rather than a settled answer. It needs the selective-prediction literature re-read on its evaluation protocol specifically rather than on its contribution, and it needs a defensible account of what "error" means when the quantity of interest is not a misclassification but an action taken on a value whose basis did not warrant it. Both are tractable, and neither is the thing that was feared, which was that the question had no answer that did not route through the designer's judgement.

It is worth recording what nearly happened here, since this page exists to show the working. The move available a moment before this one was to declare the designer's judgement an unavoidable limitation, observe that the discipline can remove human vouching from every handover except its own bootstrap, and present that as a finding rather than a gap. It is a tidy move and it reads as candour. It was also wrong, and the answer was sitting in the project's own bibliography, read months ago and filed under a different question. Declaring a problem unsolvable is a claim like any other, and it turns out to need a basis too.

What this spine is for

Published in this state, the structure can be argued with while it is still cheap to change. The parts that are firm are firm because the review forced them, not because they were chosen: the commission, the trap, the artefact's typing, the two-faced definition of failure. The part that is open is marked open. If the discrimination problem has no good answer, that is worth discovering here rather than in a chapter written as though it did.

How much of this is written up

The sections above are the plan; the continuous chapter prose is written from them one at a time and published separately, so that each can be argued with as it lands rather than held back until the whole chapter is finished. This table is the current state of that carry-over. It is part of what the spine is for: the reader can see which parts have hardened into prose and which are still only outline.

Spine section Written up as chapter prose
The commissionIntroduction and the trap
The trap, which is the same shape as the thesisIntroduction and the trap
Why design scienceIntroduction and the trap
Typing the artefactTyping the artefact
The circularity problemMarking your own homework
Criteria the design did not authorCriteria the design did not author
Success and failure, declared in advanceSuccess and failure, declared in advance
The debt the review incurredThe debt the review incurred
The discrimination problem, and a route through itThe discrimination problem, and a route through it
What this spine is forThis page; the chapter close is What is settled, what is open

References

Peffers, K., Tuunanen, T., Rothenberger, M. A. and Chatterjee, S. (2007). A Design Science Research Methodology for Information Systems Research. Journal of Management Information Systems, 24(3), pp. 45-77. doi.org/10.2753/MIS0742-1222240302