Tīrtha is an OpenAI-compatible gateway that checks an answer before it serves it. You call one endpoint and never pick a model: each request goes to the lightest model that can handle it, and a frontier model is reserved for the requests the lighter ones cannot pass. Send tool schemas and every call is checked against them. Send tests and your code runs in a sandbox, so you only get code that passes. When a request has nothing to check against, the response says so instead of calling it verified. Change one base URL and leave the rest of your client alone.
You call one endpoint and never pick a model. Behind it, the request is classified by shape and each shape gets its own path, its own check, and its own honest answer about whether anything was checked at all.
Put your checks in tirtha.tests. We run the code in a sandbox and serve only what passes. A failure escalates instead of reaching you.
Every emitted call is checked against your schema. When the schema is ambiguous it declines rather than inventing a call.
There is nothing to run, so we do not pretend. Two lightweight models must independently agree or the request escalates.
Every response carries a tirtha object beside the standard fields. Most gateways tell you nothing about how your answer was produced. This is the part we would want if we were the customer.
Each call goes to the lightest model and effort that can handle it.
The result is verified, so a wrong answer is caught, not served.
Premium models are spent on the calls that truly need them.
| Suite | Tīrtha | Notes |
|---|---|---|
| HumanEval+ (164) | Not published | Our latest HumanEval+ run has not passed our own adversarial re-check, and its saved evidence does not let us re-derive the score problem by problem, so we publish no figure for it until it is re-run. Corrected 2026-09-26: this cell previously showed a figure from that run. We make no coding accuracy claim and no comparison with frontier models here. |
| BFCL v4 tool calls | 749 / 800 | One fixed pass of BFCL v4 Python through the production tool path. An official AST re-score of the same saved outputs gives 751/800. Single pass, no multi-seed, no bootstrap interval. Not a leaderboard submission. We make no ranking claim on tools. |
| Reliability trap set, primary path | 200 / 200 | Serves no wrong tool: 0 wrong tools served across four adversarial shapes, parameterised 50 ways each (200 rows). Corrected 2026-08-06: this row previously carried a 95% confidence interval computed on the 200 rows. Rows are not the independent unit here, since the set is four shapes replicated 50 times and replicates behave as one, so that interval was withdrawn. We publish the count, not a rate. |
| Reliability trap set, failover path | 30 / 40 | When the primary stalls we fail over for availability. That model matches the primary on calling and fails it on declining: all 10 misses are the same verdict, it calls a tool where it should abstain. Measured 2026-08-02. |
| Operations | 0 retries | 0 retries in 800, 1 wrong tool in 800, deterministic routing, on the same single BFCL pass (2026-07-29). Scored on our harness and re-scored with the official checker; methods on request. |
Change your base URL to ours and keep everything else. Copy the call, drop in your key, run it in your terminal.
# change one line: your base URL curl https://api.tirtha.ai/v1/chat/completions \ -H "Authorization: Bearer $TIRTHA_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tirtha/verified", "messages": [{"role":"user","content":"Hello"}] }'
One endpoint. You do not pick a model. Send tool schemas and every call is checked against them. Send tests and your code runs in a sandbox, so you only get what passes. When a request has nothing to check against, we say so and escalate rather than guess.