Skip to content
Paula Livingstone writing · projects · tools

Attestable Evaluation

Representation Results: What the Demand Language Could and Could Not Express

This is the principal result of the evaluation, and it does not depend on execution.

artefact freezecommit + hashes domain seal39 cases, 108 atoms translation sealdemands + predictions representationgate5 of 39 eligible execution3 of 5 matched post-execution validity analysis4 validly exact, 3 of 4 each stage hashed before the next may begin; a sealed stage that changes is reported as changed the gate is what stops an approximate demand becoming a discrimination figure
Figure 5. The evaluation sequence. The representation gate is the load-bearing step: only translations classified exact reach execution, so a requirement the demand language could not express produces no verdict rather than an approximate one. The dashed box is the post-execution validity analysis, which reclassified one case and altered neither the sealed bytes nor the executed result.

Five of the thirty-nine translations were classified exact before execution. Post-execution review found that one had been mistranslated, leaving four defensible exact translations from thirty-nine frozen domain requirements.

Two accounting views

Both are reported wherever either appears. The first is what the sealed protocol produced; the second is a sensitivity analysis performed after a classification defect was discovered, and is not a replacement figure.

ViewRepresentationDiscrimination
Frozen v1, as classified and executed5 exact, 9 overconstrained, 5 underconstrained, 20 unrepresentable3/5 matched
Post-execution validity analysis4 validly exact, 10 overconstrained, 5 underconstrained, 20 unrepresentable3/4 among valid exact translations

The sealed bytes are unchanged and nothing was rerun. The reclassification is recorded as an amendment which preserves the original digests and states that no rerun occurred.

The case in question, an act whose author is entitled to judge having seen a generated ancestry, was translated with a demand excluding generated ancestry outright. A demand stricter than the warrant it translates is a conservative overconstraint by the definition fixed before translation, so the exact classification was wrong when it was made.

This does not weaken the finding. Four of thirty-nine is a smaller number than five, and any claim that the artefact can express what acts in this domain require of their evidence is falsified either way.

unrepresentable conservative overconstraint unsafe underconstraint exact 201054 39 frozen domain requirements dashed outline: the frozen classification of 5 exact, before validity review
Figure 6. Translation status across the thirty-nine frozen requirements, on the validity-adjusted view. The dashed outline marks the frozen classification of five exact translations, one of which post-execution review found to be a conservative overconstraint. Both views are reported throughout: the sealed protocol produced five exact and three of five matched, and the adjusted analysis four exact and three of four.

Failure is heterogeneous, and the layers matter

Reporting twenty unrepresentable cases as a single number would suggest twenty identical deficiencies. They are not. Representational failure occurs at four distinct layers, and the remedy differs at each.

The table allocates the twenty unrepresentable cases only. It does not account for the ten conservative overconstraints or the five unsafe underconstraints: those are translation losses of a different kind, where a demand exists and says something other than what the requirement says, and they are reported separately above.

Layer allocation of the twenty unrepresentable cases
LayerFailure meaningCases
Record modelrequired information cannot be represented at all15
Demand languagethe information exists internally but cannot be requested4
Adjudication kernelan expressible demand is evaluated incorrectly0 confirmed
Construction boundaryrequired information never enters the record1

The demand-language row is the one that most changes how the result should be read. Four corroboration cases fail because the model already holds the machinery they need and no demand can ask for it: origin disjointness is computed, an independence attestation carrying an author and a basis is required before corroboration, and a verification type carrying its own warrant exists. None is reachable through the public decision interface.

That is latent capability rather than operational representability. It remains a v1 failure, because a user cannot express the requirement through the interface the artefact offers, and a capability that cannot be requested cannot be relied upon. But it is a different kind of failure from a concept the model cannot represent at all, and it suggests a lower-complexity response.

The zero in the kernel row is worth stating explicitly: no case in this corpus showed the adjudication kernel evaluating an expressible demand incorrectly. That is a narrow finding over a small number of eligible cases and should not be read as a general correctness claim.

family stratum exact over under unrep consequence1211freshness113scope1112authority113re-establishment113corroboration15completeness1112field-provenance222semantic212ancestry2111 total 4 10 5 20 every exact translation falls in a stratum 2 family; no stratum 1 family has one
Figure 7. Translation status by scenario family, on the validity-adjusted view. The failure is systemic rather than concentrated: seven of ten families produced no exact translation at all, and corroboration produced nothing but unrepresentable cases. The pattern along the stratum column is the sharper finding. All four exact translations are stratum 2 cases, whose warrants required an inferential bridge this work supplied, and none is a stratum 1 case, whose warrant is stated outright by an external source.

Lost comparisons

Six of the nine matched pairs could not be run at all, one arm being unrepresentable and leaving no common demand to hold constant. These are lost experimental comparisons and are reported as lost rather than omitted. The pairs were the corpus's most informative shape, and losing two thirds of them is itself a consequence of the representational failure.

The most frequent missing constructs

  • a clause asserting that a named condition does not satisfy a named requirement
  • disjunction between demand clauses
  • a clause requiring an external independence attestation
  • conditional withdrawal of a demand clause
  • a representation of the act's consequence

The first is the discipline's own central move, and its absence is the sharpest form of the result. The corpus contains 24 negative atoms, and the demand language expresses that shape only where the non-entailment happens to be a property of a field, as with construction provenance, rather than a proposition about one.