Skip to content
Paula Livingstone writing · projects · tools

Attestable Methodology: First Draft (superseded)

The Discrimination Problem, and a Route Through It

The question the work is falsifiable on is whether the discipline rejects False Determinism without rejecting so much legitimate practice as to be unusable, and that is a claim about discrimination, which is only meaningful against a ground truth. The obvious sources of that ground truth fail: pre-registration fixes timing but not authorship, expert labels relocate the human vouch the work exists to remove, and incident reports document only what went wrong. The route through is a field the review already surveyed. Selective prediction evaluates systems whose job is to decline without labelling any case as one that should have been refused, using the coverage and risk relationship, and the ground truth it needs is outcome rather than opinion. One residual is real and stated: the scenario distribution stays the designer's even when the verdicts do not. The section records what nearly happened instead, which was calling the problem unsolvable and selling the gap as the finding.

Superseded. This is an earlier draft of the methodology chapter, retained as a record of what the work believed at the time. It is not current: the composition it describes includes constrained elevation, which the design chapter withdrew after implementation established that no axis of a basis is elevatable. The current chapter is the one published under the methodology category.

Everything to this point has arranged the evaluation so that it cannot be gamed by its author. The typing was read off the review, the anti-circularity commitment was inherited from the paradigm, the composition criteria were quoted from other people's specifications, and the two faces of failure were declared in advance. All of that establishes that the right thing was built and defines what failing would mean. None of it answers the question the work is actually falsifiable on, which is whether the discipline discriminates. This section is about that question, and it is the one part of the methodology that was genuinely unsolved rather than merely unwritten.

The falsifiable claim, from the problem statement, is that the discipline rejects False Determinism in realistic scenarios without rejecting so much legitimate practice as to be unusable. That is a claim about discrimination, and discrimination is only meaningful against a ground truth. Something has to say which scenarios ought to be refused and which ought to pass, or there is nothing for the discipline's verdicts to be right or wrong about. The difficulty, and it is the same difficulty the whole thesis turns on, is that the obvious sources of that ground truth put a human back exactly where the work says none should stand.

Why the obvious answers fail

Three candidates present themselves, and each fails in a way worth stating, because the route that works is shaped by their failures.

Pre-registration alone is insufficient, and the reason is visible in what the practice was designed to fix. The flexibility available in analytical choices, the researcher degrees of freedom, combined with selective reporting of the most favourable result, raises the rate of false positive and overoptimistic findings (Mandl et al., 2024), and public pre-registration is the reform aimed squarely at it, alongside data sharing and stricter reporting policies (Fletcher, 2024). Fixing the scenarios and their expected verdicts before the rules harden solves the timing half of the circularity here in the same way, since rules settled after the cases cannot be quietly fitted to them. But every one of those safeguards governs when a choice is fixed, not who fixes it, and that is the half this work cannot borrow. It does nothing about authorship. The verdicts are still the designer's honest predictions, and honest is not the same as independent. Pre-registration is necessary and it is not enough.

Expert labelling relocates the problem rather than solving it. If domain experts supply the verdicts on which scenarios should be refused, a human still vouches for the standard, which is the precise thing the review identified as where every surveyed field stops. It would move the vouch from the designer to an expert, and a vouch moved is not a vouch removed.

Incident reports cover only half the question. They document what went wrong, and in this work they are already bounded to illustration rather than evidence. Even taken at their strongest they give only the false-positive side, the cases that failed, and the harder half of discrimination is the legitimate practice that must not be refused. Failures are written up where normality is not: there is no incident report titled the plant ran correctly on Tuesday, and so no natural corpus of the cases a good discipline must let through.

The route: a field that already evaluates refusal

The way through is that this problem is not new, and the field that solved it is one the review already surveyed and read in full. Selective prediction and learning to defer build systems whose whole job is sometimes to decline, and they face the identical question: how do you evaluate a mechanism whose output is sometimes a refusal? Selective prediction gives a network an integrated reject option and evaluates it by trading coverage against the error rate on what it does answer (Geifman and El-Yaniv, 2019); learning to defer routes a case to a downstream expert on that expert's expected accuracy rather than on the model's own uncertainty (Mozannar and Sontag, 2020). Both answer the evaluation question without labelling any individual case as one that ought to have been refused, and the instrument is the coverage and risk relationship.

The instrument is simple to state. For each level of coverage, meaning the proportion of cases the system accepts rather than declines, measure the error rate over the cases it accepted. The two extreme positions need no judge at all, because they sit at the ends of the axis and are self-evidently worthless. A discipline that refuses everything sits at zero coverage, which is the problem statement's trivially safe and practically worthless expressed as a coordinate rather than a phrase. A discipline that accepts everything carries the base error rate unchanged and has added nothing. Discrimination is the shape of the curve between those two ends, and it is read off the measurements rather than adjudicated by anyone.

What this changes is the kind of ground truth the evaluation needs, and the change is what dissolves the objection. It does not need a verdict on each scenario supplied by the designer. It needs an outcome for each scenario, which is whether acting on that value was in fact warranted, and it compares the discipline's ordering of the cases against those outcomes. The distinction matters because the alternative, ranking cases by a confidence value the model emits, is ruled out by the review's finding that estimating predictive uncertainty is not sufficient for safe decision-making (Gawlikowski et al., 2021): a curve read over a signal of the wrong category would measure the discipline against precisely what it exists to replace. The designer does not label which cases should be refused. The designer builds cases whose outcomes are determinate by construction, and the measurement follows from the outcomes. The verdicts are observed, not authored, and that is the whole of the difference between this and expert labelling.

The residual, stated rather than hidden

One residual is real and is named here rather than smoothed over, because a section that claimed the objection fully dissolved would be committing the chapter's own cardinal sin. The scenario set is still the designer's. The individual verdicts are outcomes rather than opinions, but the distribution of cases is authored, and a set quietly weighted toward easy cases would flatter the curve without any single verdict being false.

Three things bear on that residual, and none of them is a claim to have removed it. The scenarios are fixed before the rules harden, so the distribution at least cannot be fitted to the discipline after the fact. The composition criteria of the earlier section test whether what was built is the missing thing at all, independently of how it scores on any distribution, so a flattering curve over a well-chosen set cannot rescue a discipline that is really a neighbour in disguise. And the distribution itself is declared and argued in the open, because a scenario set whose provenance is stated is a scenario set that can be attacked, and being attackable on that point is the honest position rather than a solved one.

What is still owed, and what nearly happened instead

This is a route rather than a settled answer, and saying so is part of keeping it honest. It needs the selective-prediction literature re-read on its evaluation protocol specifically, rather than on its contribution, which is a different reading from the one already done. And it needs a defensible account of what error means when the quantity of interest is not a misclassification but an action taken on a value whose basis did not warrant it. Both are tractable, and neither is the thing that was feared, which was that the discrimination question had no answer that did not route through the designer's judgement.

It is worth recording what nearly happened here, because this chapter exists to show its working and this is the sharpest instance of the working going one way rather than another. The move available a step before this one was to declare the designer's judgement an unavoidable limitation, observe that the discipline can still remove human vouching from every handover except its own bootstrap, and present that as the contribution rather than as a gap. It is a tidy move and it reads as candour, which is exactly why it is dangerous. It was wrong, and the answer that showed it was wrong was sitting in the project's own bibliography, read months earlier and filed under a different question. Declaring a problem unsolvable is a claim like any other, and it turns out to need a basis too. The discipline the whole work is about applied, in the end, to a claim the work was tempted to make about itself.

References

Fletcher, S. C. (2024). Reproducibility of Scientific Results. Stanford Encyclopedia of Philosophy. plato.stanford.edu/entries/scientific-reproducibility

Mandl, M. M., Becker-Pennrich, A. S., Hinske, L. C., Hoffmann, S. and Boulesteix, A.-L. (2024). Addressing Researcher Degrees of Freedom Through minP Adjustment. arXiv:2401.11537. arxiv.org/abs/2401.11537

Geifman, Y. and El-Yaniv, R. (2019). SelectiveNet: A Deep Neural Network with an Integrated Reject Option. ICML 2019. arXiv:1901.09192. arxiv.org/abs/1901.09192

Mozannar, H. and Sontag, D. (2020). Consistent Estimators for Learning to Defer to an Expert. ICML 2020. arXiv:2006.01862. arxiv.org/abs/2006.01862

Gawlikowski, J., Njieutcheu Tassi, C. R., Ali, M., Lee, J., Humt, M., Feng, J., Kruspe, A., Triebel, R., Jung, P., Roscher, R., Shahzad, M., Yang, W., Bamler, R. and Zhu, X. X. (2021). A Survey of Uncertainty in Deep Neural Networks. Artificial Intelligence Review. arXiv:2107.03342. arxiv.org/abs/2107.03342