Skip to content
Paula Livingstone writing · projects · tools

Attestable Conclusion

Contributions, and What the Evaluation Falsified

Contributions

A portable, per-claim representation of evidential basis. A record that travels with a claim, carrying how its evidence was acquired, what operation produced it, what typed events it descends from, who stands behind it, when it was established, over what it was established, and how each of those fields came to be populated. Its central commitment is negative: there is no ordering over bases, no scale on which one is stronger, and every property is stated per axis or over what a record must preserve.

Derivation-aware preservation. An inheritance rule under which what a claim rests on survives the derivations performed upon it, and under which a step cannot conceal what it was performed over. The rule is enforced in code and protected by mutation operators that reverse each correction and confirm the tests notice.

Action-relative evaluation. An admissibility method taking a basis and the demand of a specific act, returning admit, refuse or undetermined with a trace naming which axis decided and against what. The third verdict is load-bearing: a record that cannot answer what an act asks is not thereby refused, and unknown never satisfies.

An executable artefact with an executable specification. A Python library of some 2,800 lines, a specification of forty-three canonical cases of which thirty-seven exist because the theory got something wrong during design and was corrected — cases that therefore cannot serve as independent tests of it — property tests over generated input, and forty-three mutation operators all of which the suite kills. The specification was written before the implementation, so the code is tested against the published theory rather than against itself.

A negative representational finding, arrived at under a sealed protocol. The last contribution is the one the work did not set out to make, and on the evidence assembled it is the most substantial.

What the evaluation falsified

The evaluation falsified broad representational adequacy of the v1 demand language. Against thirty-nine requirements fixed beforehand in the domain's own words and grounded in external sources, five translations were classified exact before execution; post-execution validity analysis identified one mistranslation, leaving four defensible exact translations. Ten were stricter than the requirement they encoded, five were weaker, and twenty could not be encoded at all.

The pattern matters more than the count. All four defensibly exact cases were stratum 2, and no stratum 1 case was exactly representable. Stratum is a property of a case rather than of its family: it records whether the warrant is stated outright by an external source or reached through an inferential bridge this work supplied. The four occur within three families, but it is the case-level warrant that carries the finding.

Within this corpus, then, the language expresses requirements reached through this work's own reasoning and does not express requirements as external sources state them. Because the corpus is purposively allocated rather than sampled, that is a finding about these thirty-nine cases and estimates nothing about prevalence in operational practice.

The largest identified deficiency is precisely the discipline's own central move. Twenty-two of the frozen expressions state that some available fact does not itself warrant a conclusion: that a closed indication is not a proof, that separate devices are not thereby independent, that a model's account of its own working is not a check. The language can exclude a value and require one. It has no way to say that a satisfied condition fails to license a conclusion, except where that happens to be a property of a field rather than a proposition about one.

What was not falsified

Four things stand, and the distinction is not a hedge.

False Determinism is not falsified. The evaluation says nothing about whether claims acquire unwarranted authority at handover boundaries; it concerns whether the requirements that would refuse them can be written down.

Action relativity is not falsified. The demonstration and the corpus both show one basis adjudicated differently for two acts, which is the thesis working rather than failing.

The kernel is not falsified. No case showed the adjudication method evaluating an expressible demand incorrectly, though that is a narrow finding over four eligible cases and not a correctness claim.

The record model is not wholly falsified. It carried what the exactly-representable cases needed, and fifteen unrepresentable cases indicate specific absent constructs rather than a general inadequacy.

Equally, the result should not be softened into a call for further work. The demand language failed its principal representation test, on a corpus assembled and sealed before it was applied.