Skip to content
Paula Livingstone writing · projects · tools

Attestable Literature Review: The Chapter

The Fields That See the Boundary: Assurance, Governance, Explanation, and Trust

The four literatures that are aware of the boundary and still hand the judgement on. Assurance cases argue over evidence whose standing they presuppose, at design time, and name automation bias only to refer it onward. Governance mandates that basis reach the receiver and names confabulation and over-reliance as official risks, without specifying what would refuse them. Explanation approximates the model, and in deployment reaches engineers rather than the affected party. Trust in automation shows what reliance without basis costs, with fatalities in aviation and at sea, and concludes the human must calibrate. A bounded testimony coda closes it. Awareness of the boundary is not the scarce thing.

The fields surveyed so far build integrity mechanisms and stop short of admissibility because they were built to do something else. This section turns to four bodies of work that are harder cases, because each one is aware of the boundary. Safety assurance argues explicitly about whether evidence supports a claim. Artificial-intelligence governance mandates, in law and in official frameworks, that the basis of a decision reach the person who will act on it. The explainability literature exists to convey why a model produced what it produced. And the trust-in-automation field has studied, for forty years, what happens when a human relies on conveyed output without checking it. These fields do not merely stop at the boundary; they see it, name it, and in one case name the exact failure this work addresses. What none of them does is build the gate. The section closes with a bounded coda in the epistemology of testimony, which locates the handover in its deepest lineage and then departs.

Assurance cases argue over evidence whose standing they presuppose

The safety-assurance literature is the one this work must engage most carefully, because the contribution's problem statement excluded it from the comparison class, and an exclusion is only credible if the excluded literature has been read. The claim to defend is that safety engineering consumes admissibility judgements rather than producing them. A safety case, in the standard definition the field's machine-learning guidance adopts, is "a structured argument, supported by a body of evidence that provides a compelling, comprehensible and valid case that a system is safe for a given application in a given operating environment" (Hawkins et al., 2021). The argument is expressed in Goal Structuring Notation, which represents claims, strategies, context, assumptions, and evidence nodes, and organises the reasoning that connects the evidence to a top-level safety claim.

The hardest test of the exclusion is assurance applied directly to machine learning, because that is the one place the safety literature reaches into this work's territory. If any safety method produced action-relative admissibility for a model's outputs, it would be there, and it does not, for reasons the guidance states about itself. Its process "supports the development of an explicit safety case for the ML component," and the assurance activities "are performed in parallel to the development process of the ML component" (Hawkins et al., 2021). The unit is the component and its stated operating context, decided before deployment; there is no mechanism for evaluating an individual output at the moment of use. The guidance assures that a detector will identify pedestrians under a specified condition. It says nothing about whether this particular detection, right now, should be acted upon.

The decisive point is the second one, and the literature makes it against itself. The process "requires as input the system safety requirements," which are "expected to be generated by domain experts or derived from the relevant regulatory requirements," and its argument patterns then organise evidence to show "how and the extent to which the generated evidence supports the relevant ML safety claims" (Hawkins et al., 2021). Assurance structures and relates evidence whose credibility arrives from elsewhere. It adjudicates the relevance and sufficiency of evidence for a safety claim; it does not establish the epistemic standing of the evidence itself. That standing is precisely what a basis representation would carry, and it is what the assurance case presupposes. The exclusion is therefore not evasion: it is a boundary the literature draws around itself, and the whole apparatus is in any case a document that people write and reviewers read, not a function a system evaluates at a boundary.

The excluded literature then makes an admission that shapes the rest of this section. In a scoping note, the machine-learning assurance guidance observes that where a human provides oversight or fallback for a model, the safety argument must account for human-factors issues, and it names automation bias directly (Hawkins et al., 2021). This is the safety literature conceding, in a footnote, the failure this work generalises: the human receiving the output may over-trust it. It flags the risk and hands it onward to the human-factors literature rather than building a mechanism against it. That hand-off is the admissibility gap appearing inside the very field that was excluded, and it is followed to its destination below.

Governance mandates that basis reach the receiver, and does not adjudicate it

If the safety field hands the problem on, the governance field legislates it. The official frameworks locate transparency and explanation exactly at the handover, and they are unambiguous that what is owed is owed to whoever must act. High-risk systems, under European law, must be "sufficiently transparent to enable deployers to interpret a system's output and use it appropriately" (European Union, 2024). The international recommendation requires "meaningful information" sufficient "to enable those affected ... to understand the output" and "to challenge" it (OECD, 2019). The foundational risk framework makes validity and reliability the first trustworthiness characteristic and treats accountability and transparency as distinct properties to be managed (NIST, 2023). Algorithmic auditing, the field's accountability instrument, checks that a development process was carried out as it should have been, which is lifecycle accountability and not a judgement about a particular output (Raji et al., 2020).

The strongest witness is the generative-AI profile, which names this work's problem in official language. It defines confabulation as "the production of confidently stated but erroneous or false content ... by which users may be misled or deceived," and lists human-AI configuration, including automation bias and over-reliance, as a named risk category, alongside information integrity and content provenance as core considerations (NIST, 2024). The problem, the mechanism, and the remedy-adjacent concept all appear as named risks in an official framework. And yet the instruments stop where the others do. These frameworks mandate that basis reach the receiver, and they name the failure that follows when it does not; none specifies a per-output, action-relative basis representation that a receiver could refuse on. Their transparency is pitched at the model or the decision, addressed to a person. They name the boundary. They do not build the gate, and the gap this work occupies sits inside their own stated aims, unmet by their own instruments.

Explanation approximates the model; it does not evidence the basis

The natural rejoinder is that explainability supplies what governance mandates. The explainability literature's own critical strand establishes that it does not, and the argument begins before the methods do. Description and justification are different things: making a model interpretable "may not help if the goal is to assess whether the basis for decision-making is normatively defensible," and, decisively, "that decisions based on machine learning reflect the particular patterns in the training data cannot be a sufficient explanation for why a decision is made the way it is" (Selbst and Barocas, 2018). An account of how a model behaved is not a warrant that its output should be acted on, and no increment of descriptive fidelity converts one into the other.

The prevailing methods then fall short in the way that framing predicts. The dominant post-hoc techniques build simplified surrogate models that approximate the true decision criteria, which makes them scientific models in the sense that all models are wrong but some are useful, rather than explanations of a specific decision; they are "accurate representations only of a specific domain or 'slice'", they give "false assurances" when their limits are not understood, and they "do not provide evidence of the trustworthiness or acceptability of the model overall" (Mittelstadt et al., 2018). The field's own methodological survey concedes the same from the other side, treating interpretability as a means to auxiliary goals that "evaluation metrics cannot capture" (Doshi-Velez and Kim, 2017). Even the flagship legal-technical remedy declines the task: counterfactual explanations "do not attempt to convey the logic involved" and "bypass the substantial challenge of explaining the internal workings," their authors doubting whether "human-comprehensible meaningful information about the logic involved in a particular decision can ever exist," and settling instead for an actionable surrogate that lets a receiver understand, contest, and alter a decision (Wachter et al., 2018). The most-cited proposal for satisfying the transparency mandate concedes that the specific decision's basis is not recoverable and offers a useful approximation in its place.

The empirical picture completes the case. Interviews across roughly thirty organisations found that explainability as actually deployed serves internal stakeholders, the engineers debugging the model, rather than the people affected by its decisions, so that "there is thus a gap between explainability in practice and the goal of transparency, since explanations primarily serve internal stakeholders rather than external ones"; organisations "lack frameworks for deciding why they want an explanation," and gradient explanations "do not 'explain' anything to stakeholders" (Bhatt et al., 2020). The governance mandate that basis reach the receiver is, in the field, failing: explanations are consumed as engineering sanity checks and do not arrive at the handover at all. Between them, these two literatures leave the receiver with a mandate that the basis should reach them and an instrument that does not carry it.

The human must calibrate, and calibration requires the basis

The automation-bias hand-off can now be followed to the field that received it. The trust-in-automation literature has studied reliance on conveyed output for decades, and its foundational taxonomy of use, misuse, disuse, and abuse supplies the term that matters here: misuse is "overreliance on automation ... failing to monitor it effectively" (Parasuraman and Riley, 1997). That is this work's failure mode in the human-factors register, and the structural condition that produces it is the handover itself, since automation at higher levels carries out a function and informs the operator to that effect while the operator cannot control the output. The operator receives conveyed state and acts on it.

This field also supplies further high-consequence domains in which the failure is documented, which matters because a single domain invites the reply that the problem is a local quirk. The accident record includes controlled-flight-into-terrain cases "in which the crew selected the wrong guidance mode, and indications presented to the crew appeared similar to when the system was tracking the glide slope perfectly" (Parasuraman and Riley, 1997). Conveyed state was indistinguishable, at the point of reliance, from correctly founded state, and the crew acted on it. The same catalogue records a fatal case in which a pilot with low confidence in his own manual skills relied heavily on the autopilot and failed to monitor airspeed. The trust literature adds a case at sea that is purer still: the Royal Majesty ran aground because its position feed had silently reverted to dead reckoning when the satellite antenna failed, and the display went on presenting a position of unchanged appearance while its basis had quietly collapsed (Lee and See, 2004). The conveyed value was not corrupted in transit and was not obviously wrong; what had changed, invisibly, was what it rested on. The petrochemical incident of the previous section, the cockpit accidents, and the grounding at sea are three independent domains, decades and disciplines apart, documenting one failure: a receiver acting on conveyed state whose basis did not warrant the act.

Carried into the era of learned systems, the phenomenon keeps its shape and acquires sharper names. Users "tended to over-trust this advice, even in cases where it was clearly incorrect or irrelevant," a pattern the field calls automation bias or automation-induced complacency, and blind reliance is defined as "acceptance of [machine] control actions without question of its intent or motives" (Mehrotra et al., 2023). That is a receiver promoting conveyed state to fact without checking its basis. The field's goal in response is appropriate trust, and its seminal statement is careful about what trust does and does not settle: trust "guides, but does not completely determine, reliance," because trust is an attitude and reliance a behaviour, mediated by intention (Lee and See, 2004). Appropriate reliance is a correspondence between trust and the automation's actual capability, and the condition for achieving it is the load-bearing point for this review: calibrating trust requires understanding the automation's purpose, process, and performance (Lee and See, 2004) (Mehrotra et al., 2023). Appropriate reliance requires knowing why and how, which is to say it requires the basis, which is exactly what an unattested claim strips away.

One limitation bounds this. The classic human-factors work predates generative models: automation bias was established for alarms, autopilots, and decision aids, not for systems that compose fluent assertions. Applying it to machine-generated claims is a reasoned extension this work argues, with the systematic review as the bridge, rather than a finding lifted whole.

The thread that began in the excluded safety literature now closes. Safety engineering named automation bias and handed it to human factors. Human factors has studied it for forty years and concluded that the remedy is reliance calibrated to the automation's purpose, process, and performance. Both fields therefore arrive at the same gap from opposite sides, and both stop in the same place: at the human who must calibrate. Neither builds the mechanism that carries the basis across the boundary. That omission is defensible while the receiver is a person who can, in principle, ask. It is not defensible at a boundary where the receiver is a machine that cannot ask, and has no basis to calibrate against.

A coda: the handover has an older name

The structure this chapter has traced, one party taking another's say-so and acting on it, is the oldest problem in the epistemology of testimony, and the review touches it once, for depth, and departs. Two points earn their place. The first is that the field's central dispute is about this exact boundary: the transmission view holds that testimonial knowledge can only be passed on, so that a hearer cannot acquire what the speaker lacked, while the generation view holds that testimony can create justified belief even where the speaker's own evidence did not justify the speaker (Leonard, 2021). Whether authority can increase across the speaker-hearer boundary is precisely the question a gate at that boundary exists to answer.

The second is that philosophy has the case this work is built around, and does not have a name for it. The persistent-believer case describes a hearer justified in believing a claim on a speaker's say-so although the speaker's total evidence defeated the speaker's own justification (Leonard, 2021). That is the promotion this work names: the receiver ends up holding a claim as better warranted than its basis at source supports. The named views in the vicinity, inheritance and assurance, describe how trust passes rather than how authority is elevated beyond its basis.

The departure is as important as the engagement. Epistemology asks whether the hearer is justified. This work asks something narrower and mechanical: whether a system can carry the basis across the boundary so that the elevation can be refused by something other than the hearer's judgement. Philosophy diagnoses the boundary and debates the hearer's warrant; it does not build a mechanism that marks basis and rejects unwarranted promotion at a machine handover, where there is no hearer to exercise judgement at all. That is the review's business, and the chapter returns to it.

What this section establishes

These four fields are the ones that see the boundary, and the pattern across them is sharper than in the fields that do not. Safety assurance argues rigorously about evidence and presupposes its standing, then names automation bias and hands it on. Governance mandates that basis reach the receiver and names confabulation and over-reliance as official risks, without specifying anything a receiver could refuse on. Explainability is the instrument that mandate reaches for, and its own critical literature establishes that it approximates the model rather than evidencing a decision's basis, and that in practice it does not reach the affected receiver at all. Trust in automation demonstrates, with fatalities in the cockpit and at sea, what reliance without basis costs, and concludes that reliance must be calibrated to purpose, process, and performance, which is to say to a basis that nothing in the stack carries. Awareness of the boundary, it turns out, is not the scarce thing. Every one of these fields ends by handing the judgement to a human, and the handover this work is concerned with is the one where no human is standing.

References

Hawkins, R., Paterson, C., Picardi, C., Jia, Y., Calinescu, R. and Habli, I. (2021). Guidance on the Assurance of Machine Learning in Autonomous Systems (AMLAS). Assuring Autonomy International Programme, University of York. arXiv:2102.01564. arxiv.org/abs/2102.01564

NIST (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. doi.org/10.6028/NIST.AI.100-1

European Union (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 13: Transparency and provision of information to deployers. Official Journal of the European Union. eur-lex.europa.eu/eli/reg/2024/1689/oj

OECD (2019). Recommendation of the Council on Artificial Intelligence, Principle 1.3: Transparency and explainability. OECD/LEGAL/0449. legalinstruments.oecd.org

NIST (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. doi.org/10.6028/NIST.AI.600-1

Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D. and Barnes, P. (2020). Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing. FAT* 2020. arXiv:2001.00973. arxiv.org/abs/2001.00973

Selbst, A. D. and Barocas, S. (2018). The Intuitive Appeal of Explainable Machines. Fordham Law Review, 87(3), pp. 1085-1139. ir.lawnet.fordham.edu/flr/vol87/iss3/11

Doshi-Velez, F. and Kim, B. (2017). Towards A Rigorous Science of Interpretable Machine Learning. arXiv:1702.08608. arxiv.org/abs/1702.08608

Mittelstadt, B., Russell, C. and Wachter, S. (2018). Explaining Explanations in AI. FAT* 2019. arXiv:1811.01439. arxiv.org/abs/1811.01439

Bhatt, U., Xiang, A., Sharma, S., Weller, A., Taly, A., Jia, Y., Ghosh, J., Puri, R., Moura, J. M. F. and Eckersley, P. (2020). Explainable Machine Learning in Deployment. FAT* 2020. arXiv:1909.06342. arxiv.org/abs/1909.06342

Wachter, S., Mittelstadt, B. and Russell, C. (2018). Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR. Harvard Journal of Law & Technology, 31(2), pp. 841-887. arXiv:1711.00399. arxiv.org/abs/1711.00399

Parasuraman, R. and Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors, 39(2), pp. 230-253. doi.org/10.1518/001872097778543886

Lee, J. D. and See, K. A. (2004). Trust in Automation: Designing for Appropriate Reliance. Human Factors, 46(1), pp. 50-80. doi.org/10.1518/hfes.46.1.50_30392

Mehrotra, S., Degachi, C., Vereschak, O., Jonker, C. M. and Tielman, M. L. (2023). A Systematic Review on Fostering Appropriate Trust in Human-AI Interaction. arXiv:2311.06305. arxiv.org/abs/2311.06305

Leonard, N. (2021). Epistemological Problems of Testimony. Stanford Encyclopedia of Philosophy. plato.stanford.edu/entries/testimony-episprob