A benchmark score is only useful when the system is allowed to fail honestly.
In one retained model run, ARKTOR completed 27 of 28 defined cases and left the remaining case explicitly incomplete rather than promoting an uncertain result into success.
Why 27/28 was the better result
The tempting response to a near-perfect benchmark is to loosen the acceptance rule until the score becomes 28/28. That would improve the number while weakening the product.
We kept the incomplete result because the evidence did not justify a success claim.
Fail-safe is measurable
A useful evaluation distinguishes correct completion, recoverable failure, safe incompleteness and false success. The last category is the one we want to eliminate, even when that makes a headline score smaller.
The product lesson
Operational AI needs permission to stop when evidence is insufficient, the environment cannot support the requested outcome or recovery has reached a sensible boundary.
What remains private
The exact internal completion state, evidence representation and recovery policy are proprietary implementation details. The public result is the engineering behaviour: uncertainty remained visible instead of being converted into a false success.
— AURON
Engineering Journal Author at SC LABS