Skip to content
Paula Livingstone writing · projects · tools

Attestable Evaluation

Implications: A Classified Failure Set, and Why No Repair Was Attempted

The classification was made after execution and before any design work, to establish whether one general relation would answer these failures or whether several narrower constructs are needed.

The unit needs stating precisely, because the chapter has moved between atoms, expressions, cases and translation statuses. Twenty-two is the number of frozen requirement expressions containing a non-entailment relation. Those twenty-two expressions contain twenty-four negative atoms between them, two cases carrying two each. It is not a count of mutually exclusive failed translations.

The categories overlap. Twenty-five assignments fall across twenty-two distinct cases, and three cases appear in two categories each. The counts below are therefore per category and are deliberately not summed. The overlaps column names each shared relationship from both sides, so the three doubled cases appear there as three category pairs — fixed non-entailment with independence, fixed non-entailment with conditional, and conditional with epistemic — not as more.

Classification of the twenty-two expressions containing a non-entailment relation
CategoryCasesRepresentative identifiers (a sample; not exhaustive)Layer implicatedOverlaps
Fixed non-entailment rules8S-consequence-1, S-scope-1, S-scope-3, S-reestablishment-1demand languageconditional, independence
Missing verification requirements3S-semantic-1, S-semantic-3, S-authority-4demand language, machinery presentnone
Independence and licensing rules4S-corroboration-1, S-corroboration-2, S-corroboration-3, S-corroboration-5demand language, machinery presentfixed non-entailment
Conditional and exception logic4S-freshness-1, S-completeness-1, S-completeness-3demand language; part outside any demand languagefixed non-entailment, epistemic
Epistemic-state distinctions5S-authority-1, S-authority-2, S-freshness-3, S-freshness-4record model and demand languageconditional
Construction-boundary observability1S-ancestry-3construction boundary; no language construct appliesnone

Fixed non-entailment rules. A stated fact is true and does not license the conclusion it appears to license. A closed indication is not a proof; documentation agreeing with itself is not documentation agreeing with the plant. The pairs are enumerable rather than open, which suggests a declarative clause naming a condition and the requirement it fails to satisfy, checkable by comparison rather than by inference.

Missing verification requirements. Not entailment at all. The requirement is that a check was performed by something other than the thing being checked, and recorded. The model already carries a verification type with its own author and basis.

Independence and licensing rules. Whether two supports may be counted as two. The model holds a precise position and the demand language cannot ask for it.

Conditional and exception logic. Disjunction between clauses and conditional applicability. Part of this category is outside any demand language's scope: extending an isolation boundary and deferring work are actions, and a language describing what a basis must be cannot express what to do when it is not.

Epistemic-state distinctions. Not knowing as a third state, distinct from knowing-not. The kernel expresses these as outcomes through its three-valued verdict; a demand cannot state them as requirements. One case needs the record to distinguish a recorded negative from an absence, which is a representation change rather than a demand change.

Construction-boundary observability and trust. One case, and no language construct addresses it. This category exists so that the observability failure is not misdiagnosed as a demand-language problem.

What the classification suggests

One general entailment relation does not appear justified. The largest group needs machinery the model already holds to be reachable from the decision interface, which identifies a plausible lower-complexity response rather than a demonstrated one: whether such exposure would yield exact translations can only be established by translating against a revised language, and it is not a trivial interface change. Deciding what a demand may ask about a basis is a semantic commitment, and exposing internal state through a public interface carries security consequences that would need their own analysis.

Why no repair was attempted here

An inference-policy layer would change the artefact's semantic category. The present language asks whether a record possesses required properties; a non-entailment relation asks whether one proposition licenses another. That is a different kind of artefact and would require a new formal account, a new specification and mutation set, a new translation, a new freeze, and a genuinely fresh corpus.

These thirty-nine cases could then demonstrate regression improvement, showing that identified losses are now represented. They could not confirm the revised artefact, having shaped it. Confirmation would need an independently authored holdout corpus.

Repairing the language within this work would also have erased the result. The finding is that this artefact, as specified and frozen, could exactly represent four of thirty-nine externally grounded requirements. Changing the artefact and rerunning would replace a measured boundary with an unmeasured one.