Benchmarks

Tested at the frontier.
It holds.

We put our models through public, independent tests designed to break them.

Holds up on new data
TableShift, public benchmark

Domain shift is when new data stops matching what a model learned on, and it is where most models degrade. With an entire region held out of training, the Clone's 90% range still covered 90% of outcomes in that new region, the same as in-domain.

Honest at the edge
Coverage and abstain rate

When an input falls outside what a Clone has seen, it abstains instead of guessing, and it says so. Coverage and the abstain rate are measured on held-out cases and travel with every prediction. Honesty is the point.

Accurate
TabArena, public benchmark

Accuracy is how close a prediction lands to the truth. On a public test a single model tracks the real answer 94% of the way (R2 0.94), within about 4% of the top published result, and it returns a calibrated range with every prediction.

Fast and efficient
MLCommons LoadGen, public benchmark

Speed is how efficiently a model runs live. On the industry-standard harness the Clone served tens of thousands of predictions per second, on an ordinary computer with no GPU. Each model is kilobytes, small enough to run anywhere.

We are building our own benchmark
for the trust axis today's leaderboards do not measure:
whether a model stays honest when the conditions change.