Notes

Testing evidence lineage against real versioned scientific records

The evidence lineage layer makes a specific claim: when a source record is
revised, it can separate the downstream work that is genuinely affected from
the work that remains valid, replay the affected parts that are reproducible,
and route the rest to a qualified human.

That claim is only worth anything if it is tested against evidence that
actually changes, chosen by someone other than us.

The test

Public scientific databases are unusually good ground for this. They are
versioned, the revisions are official and dated, and the transitions between
versions are a matter of public record rather than something a vendor gets to
define after the fact.

Internal held-out testing detected 100 of 100 selected official major evidence
transitions
, with repeatable execution across runs.

What that result is not

It is an internal research result. It is not independent validation, and it
does not become independent by being repeated or by being described more
confidently. Three specific limits apply:

  1. We selected the cases. A held-out set chosen by the party being tested is
    weaker evidence than a set chosen by someone else, however carefully it is
    constructed.
  2. Detection is the easier half. The commercially important measure is
    affected versus preserved accuracy, in both directions. A system that flags
    everything scores perfectly on detection and is worthless.
  3. The definition is ours. What counts as a material change was defined by
    us. That definition is the load-bearing part of the entire claim.

The definition problem

That third point deserves more weight than the result itself.

Deciding that a change is material and another is not is a judgement encoded as
a rule. Encode it too loosely and everything downstream is flagged. Encode it
too tightly and something that mattered slips through. Either way the number at
the end reflects the definition at the start, and right now that definition has
only been agreed by the people who wrote it.

So the next step is not a bigger internal number. It is having a domain
scientist independent of this work agree the definition in writing, before
any external benchmark is run. Until that happens, the honest position is that
we have a repeatable method and an internal result, not a validated one.

Why this is published in this form

A result with its limits attached is more useful than a result without them,
because the limits tell you what would have to be true for the result to
survive contact with someone else. If a number cannot be stated alongside what
would falsify it, it is marketing.

The full claims register, with every claim classified by how strong the
evidence actually is, is on the evidence page.

Start with one workflow

Thirty minutes on a workflow where a change in rules or evidence has already cost you work.

Next · See it work AI Control Tower