Evidence

Every claim, its class, and how to check it

Most companies ask you to trust their numbers. This page lists ours, classifies each one by how strong the evidence actually is, and tells you how to verify it without our help. It also lists what we withdrew, and why.

The standard

Five classes, and the rule that governs them

A claim that cannot be checked by someone outside this company does not belong in the top two classes. Applying that rule honestly is what removed several things we would have preferred to keep.

Verifiable now

You can check this yourself, today, without our help and without our source code.

Internal result

Measured by us, reproducible by us, not yet checked by anyone else. Never described as validation.

Partner-verified

Tested by a third party who has integrated the system. Stronger than our own result, weaker than independent, because the tester has an interest in the integration working.

Relationship

Something that has actually happened between us and a third party, stated with its exact limits.

Untested assumption

Load-bearing for a decision but not yet proved. Listed so nobody mistakes it for a result.

Withdrawn

We published it, found it did not hold, and retracted it. Left visible on purpose.

How to read the proof

Claim → status → evidence → verifier → limitation

The useful question is not whether a number looks impressive. It is whether a reviewer can follow the claim to a record, verify the record independently and see exactly where the claim stops.

01 · CLAIM

What is being said?

Start with the exact capability or measured result, not a marketing category.

02 · STATUS

How strong is it?

EVIDENCED, MEASURED, PILOT, DESIGN ONLY or NOT CLAIMED.

03 · EVIDENCE

What record supports it?

Open the programme row, signed pack, run result or published limitation.

04 · VERIFY

Can you check it?

Use the machine-readable result, standalone verifier or public gate.

05 · LIMIT

Where does it stop?

Read the boundary before treating a result as a production conclusion.

06 · DECIDE

What happens next?

A bounded pilot turns the open question into customer-owned evidence.

Evidence by programme

Where each programme's evidence actually stops

Evidence produced on a development host cannot be labelled as evidence from a representative environment, and representative evidence cannot be labelled as target evidence. The build refuses to emit the higher classification when the conditions for it are absent. Where a target tier is empty below, it is empty because the work has not been done.

Enterprise assurance

TARGETREP_AWS

The target-representative AWS run is now recorded alongside the local capability register. It is target-environment evidence, not production customer validation or certification.

Frozen regression suite
Evidenced
Runs on every build. A failure blocks the baseline seal.
Capability register
Partial
28 of 34 capabilities evidenced. The remaining six are listed with their blockers rather than rounded up.
Signing-key custody
Evidenced
AWS KMS ECC_NIST_P256 was configured and independently checked through GetPublicKey.
Target-representative cloud
Evidenced
AWS ap-southeast-2: 117,290 completed verifications, 0 performance errors, 128.1 verifications/sec soak and clean Terraform teardown.

Life sciences

PARTNER_INTEGRATION

The strongest external evidence we hold, and still an integration audit by the integrating party rather than an independent one.

Partner integration suite
Evidenced
126 integration tests passing inside a third party's platform. Their suite, their environment, our engine.
Adversarial correctness
Evidenced
Zero gate bypasses across 395 adversarial inputs. Correctness-adversarial, not load-adversarial, and the distinction matters.
Evidence-change detection
Partial
100 of 100 on a set we selected. An internal research result, not validation.
Regulatory certification
Not claimed
None held. Controls are built to be auditable, which is a different claim.

Mission-critical systems

MISSION_CRITICAL_EDGE_REP

The newest programme and the one with the least evidence. The fault-injection benchmark is the honest weak point and is described as such on the programme page.

Unit and attestation suite
Evidenced
83 tests passing, including the bitflip sweep, signature transplant and attacker-key forgery cases.
Fault-injection benchmark
Under remediation
OrbitGate prototype benchmark, under active validation. Corpus reproduction 101 of 160, up from 54. Frozen at v0.x rather than tuned further, because the rules were developed with the labels visible. Not used as evidence of aerospace readiness. The next result comes from a blind, partly partner-defined fault set.
Representative edge hardware
Pending
No run exists on this tier.
Aerospace hardware-in-the-loop
Pending
Requires a partner environment.
Radiation qualification
Not performed
Not performed, and outside what software testing can establish.

There is no total on this page, on purpose

An aggregate figure lets the strongest programme carry the weakest. Read the rows. The mission-critical fault-injection row is the weakest thing we publish, and it is published rather than omitted because a benchmark that fails the first time an engineer runs it costs more than having no benchmark at all. The programme page states it in full.

Claims register

All 23 public claims, in one table

If a claim we make anywhere else is not in this table, it should not have been made. Tell us and we will either evidence it or remove it.

23 claims
ClaimClassHow you check it
Verification is deterministic. The same input produces the same signed verdict. Verifiable now Run the same input through a public gate twice and compare the certificates byte for byte.
Every verdict is Ed25519 signed and anchored in a tamper-evident chain. Verifiable now Call the live endpoint below and verify the signature in your own code. The public key is published.
ProofBench-X: 1,116 exact-arithmetic cases, all passing. Verifiable now Clone dtl-mathgate and run the suite. It is the benchmark we do stand behind.
The gates are independently auditable without our source code. Verifiable now The audit protocol is public. Any lab can test a gate and sign a verdict in the registry.
Security lanes cover ten CWE and OWASP weakness families with zero false positives against an unguarded baseline. Verifiable now Clone dtl-security-benchmark and run it. The corpus and the scoring are both in the repository.
Our own fuzzer bypassed a naive canonicalisation defence with three obfuscations, and the hardened lane then held. Internal result Reproducible from the benchmark repository. This is our strongest security result and it is a story about finding our own weakness.
DELA detected 100 of 100 selected official evidence transitions in held-out testing, with repeatable execution. Internal result Internal research result on a set we selected. Not independent validation, and it does not become independent by being repeated.
126 integration tests pass at 100%, covering the verification layer inside a third party's platform. Partner-verified From the partner's audit report, which we can share under NDA. Their test suite, their environment, our engine.
Zero wrong answers across 1,866 correctness checks, and zero gate bypasses across 395 adversarial inputs spanning ten OWASP weakness families. Partner-verified Same report. Correctness-adversarial, not load-adversarial, and the distinction matters.
Transport failures are recorded as degraded rather than scored as valid safety abstentions. Partner-verified The partner's audit examined this specifically and hardened it. An outage cannot disguise itself as the system declining to answer.
The evidence chain survives concurrent writes from multiple processes with no duplicate or missing sequence numbers. Verifiable now Run the concurrency harness in our repository. It re-verifies the whole chain afterwards.
Metering lost charges under concurrent workers before we fixed it. Withdrawn Found by our own harness: 40 lost charges across 480 concurrent calls. Fixed with row-level and cross-process locking, and the harness now passes. Listed here because the defect was real.
A European compliance platform completed a technical review of the verification layer and integrated a first version for testing. Relationship A mutual NDA is executed. There is no commercial contract, no revenue, and no signed memorandum.
A proposal for an evidence-lineage shadow pilot was forwarded internally at a major AI research organisation for consideration. Relationship That is the entire status. No validation, adoption, testing or partnership has occurred, and we are awaiting any further response.
Four Australian provisional patent applications are filed covering DTL, DAX, DCLA and DELA. Relationship Filed, not granted. No freedom-to-operate search has been run. Our own prior art note concludes the individual mechanisms are not novel and the open question is the combination.
Regulated organisations will pay for deterministic evidence. Untested assumption Untested. EcoKure is pre-revenue. This is the single assumption the whole business rests on and we will not dress it up as anything else.
Implementation effort per customer is low enough to support a product margin. Untested assumption Unmeasured. We measure it during the first paid pilot rather than quoting a number we cannot support.
The orbital fault-injection benchmark reproduces 101 of its own 160 expected outcomes, and 27 of the 78 cases that should block still do not block. Withdrawn Clone orbitgate and run the corpus yourself. We publish the shortfall rather than the corpus size, because the corpus size is not a result. No orbital benchmark figure should be cited by us or anyone else until this reproduces in full.
Four lanes of that benchmark now reproduce fully: command 10 of 10 and propagation 20 of 20, both from 0 and 1 respectively, with 88 new regression tests. Verifiable now Clone orbitgate and run the test suite: 171 tests pass, up from 83. The new tests fail against the previous commit and pass against the current one, which is the only useful way to state a fix.
Tamper detection on signed decision records: a bitflip sweep across the entire signed envelope, signature transplant, and attacker-key forgery are all rejected. Verifiable now Clone orbitgate and run the attestation suite. The test performs the attack rather than describing it, and 83 tests pass on the current code.
Earlier internal coding-benchmark scores. Withdrawn Withdrawn in full. See the retraction record below. We publish no score for that benchmark.
An earlier 500 of 500 result on a public software benchmark. Withdrawn Withdrawn after an audit found answer-key lanes active in the run. Never cited by us, and it should not be cited by anyone else.
A 162 of 166 security-benchmark figure carried in earlier decks. Withdrawn Withdrawn. It did not reproduce from the current harness and traced only to a self-audit document. The substance was fine. The number was not.
Verify it yourself

Three ways, none of which involve us

01

Call the live endpoint

No account, no key. What comes back is a cryptographically signed verdict with its position in the tamper-evident chain.

ecokure.net/api/safegate/check?url=example.com
02

Clone a gate and break it

18 gates are public repositories. Run one against your own cases and find where it is wrong. We would rather you found it than a customer did.

Browse the gates

03

Sign your own verdict

The audit protocol is published. Any lab can test a gate and publish a signed result, without our source and without telling us first.

Audit registry

Retraction record

Three results we published and then withdrew

This section exists because deleting mistakes is how trust dies. Each entry states what the claim was, what we found when we checked it, and what we did. The original posts stay online beneath their retraction notices.

Orbital verification benchmark, 160 cases

August 2026
What we found
Re-running the corpus reproduced 54 of its 160 expected outcomes. Of the 78 cases that should have been blocked, 58 were not, including command claims that read as instructions to disable autonomous fault protection. The stored certificate no longer reproduced from the current code.
What we did
Every orbital benchmark figure withdrawn. Remediation is under way lane by lane: the corpus now reproduces 101 of 160 and unsafe non-blocks are down to 27 of 78. The remaining shortfall stays published and the benchmark stays under remediation until it reproduces in full.

Public software-benchmark result, 500 of 500

July 2026
What we found
An internal audit found answer-key lanes were active during the run. The harness could see information it should never have had.
What we did
Withdrawn in full. The original post stays online, unedited, beneath a retraction notice.

Security-benchmark figure, 162 of 166

July 2026
What we found
Re-running the suite did not reproduce the figure. It traced to a self-audit document rather than to the harness, and the arithmetic in the source did not match.
What we did
Withdrawn from all materials. Replaced with results that reproduce from the current code.

Internal coding-benchmark scores

July 2026
What we found
The test harness could be gamed. We then built a scanner to catch that class of cheating, and the scanner itself had blind spots. Twice. Two later results that passed the scanner did not survive hand inspection.
What we did
Every figure withdrawn. We publish no score for this benchmark, because there is no public submission portal and therefore no way for you to check one.

Why this is on the site rather than buried

A benchmark figure you cannot reproduce is worth less than no figure at all, because it fails the moment somebody tries. Publishing the retractions costs us three impressive numbers and buys the only thing that matters in a regulated sale: that the numbers we do publish can be trusted. If you are evaluating us, this section should carry more weight than any result above it.

Open questions

What we have not answered yet

These are the questions a serious evaluator would ask. We would rather you read them here than assemble them yourself in week three of diligence.

  • The orbital fault-injection benchmark does not currently meet its own expected outcomes, and no mission-critical evidence exists on representative embedded hardware.
  • The claim-classification layer does not generalise beyond the phrasings it was developed against: a sealed holdout returned 16 of 80 unsafe cases propagated, 0 of 40 legitimate commands passed, and no engagement at all on 56% of cases.
  • No freedom-to-operate search has been run on any of the four provisional applications.
  • No independent party has yet agreed the definition of a material evidence change, which is the load-bearing definition in the whole evidence-lineage claim.
  • No customer has taken a pilot through to a production licence, so willingness to pay is unproven.
  • Implementation effort per customer has not been measured, so gross margin is an assumption.
  • No third-party security or technical assurance report exists yet. The partner audit is an integration audit by the integrating party, which is a different thing.
  • The concurrency fix is proven on the file store. The Postgres path uses row-level locking and has not yet been tested under load.
  • EcoKure holds no certifications, and every control claim should be read as audit-ready rather than certified.

Found something wrong on this page?

That is the most useful message you could send us. A claim that does not hold should be corrected or withdrawn, and we would rather hear it from you now than from a customer's counsel later.

Next · See where it runs Deployment