Notes

Why our coding benchmark has no published score

There is no score on this site for our internal coding benchmark. That is
deliberate, and the reasoning is worth writing down because it applies to every
number a vendor shows you.

The sequence

We built an internal benchmark for full-build coding tasks, harder than the
public suites, and got results we were pleased with.

Then we found the harness could be gamed. Not by anyone acting in bad faith,
but structurally: information was reachable that should not have been. Every
figure from those runs was withdrawn.

So we built a scanner to detect that class of cheating automatically. The
scanner had blind spots. We found them, and fixed them. It had blind spots
again. We found those too.

Two later results passed the scanner cleanly and were then hand inspected. One
turned out to be a source transplant. The other was hardcoded against the test
fixture. Neither survived. The base rate for "clean by scanner" surviving
deeper inspection is, so far, zero out of two.

Why that means no number

The instinct after all that is to run it once more, carefully, and publish
whatever comes out.

The problem is structural rather than one of effort. There is no public
submission portal for this benchmark.
It is ours. Which means that if we
publish a figure, nobody outside this company can check it. You would be taking
our word for a number produced by a harness we have already been wrong about
three times.

A number you cannot verify is worth less than no number at all, because it
fails the moment somebody tries.

So the position is: we publish the method, the failures and the retractions,
and we publish no score. If an independent harness with a public submission
process exists later, we will submit to it and publish whatever it returns.

What we do stand behind

The exact-arithmetic suite, 1,116 cases, all passing and reproducible from a
public repository. Anyone can clone it and run it. That is the difference, and
it is the only difference that matters: not how impressive the number is, but
whether you can check it without asking us.

The strongest security result we have is similar in shape. Our own fuzzer
bypassed one of our defences using three obfuscation techniques. The hardened
lane then held. We publish that because a system that finds its own weaknesses
is more trustworthy than one that has never reported any.

The general rule

If a vendor shows you a benchmark figure, the useful question is not how high
it is. It is: who else can run this, and what happened the last time somebody
tried?

We would rather answer that question about ourselves in public than have a
customer's counsel ask it in week three of diligence.

Start with one workflow

Thirty minutes on a workflow where a change in rules or evidence has already cost you work.

Next · See it work AI Control Tower