What is true right now.

Most vendor numbers are a PDF from a launch date. This page is generated from the configuration we are actually running, so it cannot quietly outlive it.

For every path through the gateway: what good means on that path, whether a real oracle checked it or nothing did, what we measured, and whether that measurement still describes the system serving your request. When it does not, we strike it and say so.

Generated 2026-08-02 18:28 UTC from the live gateway configuration and our claims ledger. Every badge below is derived from that configuration, not written by hand.

The paths

Tool calls

Verified

Oracle: Real. The emitted call is checked against the tool schema.

Emits a schema-valid call, and declines rather than inventing one when the schema is ambiguous.

749 of 800 on a single fixed pass of BFCL v4 Python, through the production tool path

An official AST re-score of the same saved outputs gives 751 of 800. Single pass: no multi-seed and no bootstrap interval. Wrong tool 1 of 800. Retries 0 of 800. These are 800 fixed benchmark items replayed through the production path, so it is an artifact measurement and not an operational number from organic traffic. Not a leaderboard submission.

Adversarially reviewed. Adversarially reviewed 2026-07-30; publishable in this n-of-N single-pass form only.

config premise holds against live config (ledger L27) · adversarially reviewed 2026-07-30; publishable in this n-of-N single-pass form only

Watch it decline a trap →

Tool calls, adversarial

Verified

Oracle: Real. A wrong tool served is a hard, countable failure.

Serves no wrong tool when the request is built to induce one.

200 of 200 on the adversarial trap set, replayed through production

Clopper-Pearson 95% interval [98.1, 100], and wrong answers served 0 of 200, Clopper-Pearson [0.0, 1.9]. That interval is Clopper-Pearson, not the rule of three, which would give 1.5 rather than 1.9. The 200 traps are a strict superset of the historical 40, extended by appending and verified byte-identical for the first ten.

This is a primary-path result. See Failover below for where it does not hold.

Adversarially reviewed. Adversarially reviewed 2026-07-30; the interval label was its objection and is applied here.

ledger row L24 declares no PREMISE · adversarially reviewed 2026-07-30; the interval label was its objection and is applied here

Run the trap set →

Code, with your tests

Current

Oracle: Real. Your tests are executed in a sandbox and must pass.

Only code that passes the tests you supplied is served.

92.1% HumanEval+, cache-free (0 cache serves)

Frontier parity, not a claim of beating it (Sonnet 92.7, Opus 93.3 on our harness), and +11pp over the bare lightweight model.

config premise holds against live config (ledger L30)

Send your own tests →

Free prose

Verified

Oracle: NONE. There is nothing to run.

We do not pretend agreement is proof. Two different lightweight models must independently agree or the request escalates. Agreement is a confidence signal.

Agreement is not verification, and we label it that way

On HaluEval faithfulness judgements, measured cross-model on the pair now in production and recomputed from raw rows on 2026-07-28, the gate accepted a wrong answer in 11 of 40 agreements. That is 27.5%, Wilson 95% interval 16.1 to 42.8. The sub-splits are too small to quote as separate rates. We do not have comparably audited figures on other kinds of request, so do not export this number to them, and do not read it as a product-wide error rate: it is conditional on one benchmark's agreed slice.

Adversarially reviewed. Adversarially reviewed 2026-08-02, which rejected the previous draft; this is the wording it prescribed.

config premise holds against live config (ledger L78) · adversarially reviewed 2026-08-02, which rejected the previous draft; this is the wording it prescribed

Agentic multi-turn

Verified

Oracle: Real. Berkeley's own stateful checker decides whether the end state is correct.

Holds a multi-step task together across turns and leaves the world in the right state.

106 of 199 on BFCL multi-turn, scored by Berkeley's stateful checker, through live production

53.3%, Wilson 95% interval 46.3 to 60.1, with 0 errors. That is the same band as o1-2024-12-17 FC at 53.00% and Claude-3.7-Sonnet FC at 54.50% on Berkeley's June 2025 snapshot. It is not a leaderboard rank. Before any score was printed the instrument was validated in both directions: a positive control passed 199 of 199, and a negative control of silent and sabotaged transcripts was accepted 0 of 199. We run 199 of the 200 items, excluding one that crashes Berkeley's own checker.

Adversarially reviewed. Cleared 2026-08-02 with its comparability caveats attached.

ledger row L21 declares no PREMISE · cleared 2026-08-02 with its comparability caveats attached

Failover

Verified

Oracle: Availability, not identical safety.

Honestly: the failover model matches the primary on calling and fails it on abstaining.

Run exactly as production pins it, the failover model scores 30 of 40 on the traps

All 10 misses are the same verdict: it calls a tool where it should decline. It matches the primary on calling and fails it on abstaining, which is precisely what the trap set exists to measure. Wilson 95% interval 59.8 to 85.8; at n=40 that is wide, so do not read 30 of 40 as a precise rate. On standard tasks the same model scores 382 of 400. When the primary stalls we fail over for availability, and on that path the guarantee above does not hold. We would rather you knew.

Adversarially reviewed. Cleared 2026-08-02 to be published as the counter-example beside the primary-path number, never alone.

config premise holds against live config (ledger L20) · cleared 2026-08-02 to be published as the counter-example beside the primary-path number, never alone

Answer reuse

Measured zero

Oracle: Reuse of previously checked answers only.

Nothing is reused that was not checked when it was first served.

Zero external customers have ever been served a reused answer

65 reuse serves in 93,242 requests, all internal. Cross-context reuse only fires when the caller sends tests, and no external key has. We do not sell a cost curve on this.

Adversarially reviewed. Re-derived from the production database 2026-08-02.

ledger row L22 declares no PREMISE · re-derived from the production database 2026-08-02

What is serving, right now

You call one model id. Behind it, the request is classified by shape and each shape has its own path. These values are read from the running gateway configuration when this page is generated, which is what lets the badges above mean anything.

We publish results, not the recipe, and we would rather say so than look coy about it. Which specific model serves which request shape is the output of our own evaluation work, and it is the part a competitor could copy in an afternoon. So we show that each value was read live and that the paths genuinely differ, and we hold the map itself. Everything the map produces, including where it performs worse, is on this page. If you need to reproduce a number, reproduce it against the gateway: that is the thing we are actually selling, and it is the thing our numbers describe.

Tool-call pathconfigured, read live at generation time
Code and general pathconfigured, read live at generation time
Witness / failoverconfigured, read live at generation time
Escalation targetconfigured, read live at generation time
Escalation wall-clock budget45

What we have not measured

Stated because its absence changes how you should read everything above.

Every number here carries its population, sample size and date, or it does not appear. Claims we have withdrawn are listed on Research. This page is regenerated on deploy; if the configuration changes and the page is not regenerated, the build fails rather than publishing a stale claim.