Tīrtha is the team your developers' tools already call. The everyday majority is handled by lighter models for a fraction of frontier prices; only the genuinely hard slice pays for a frontier model. Frontier models are used only for the genuinely hard calls, so a frontier budget goes materially further. How much further depends on the workload, and we do not quote a multiple. On our own monitor and benchmark traffic, which is tool-calling heavy, blended cost measured $0.000149 per request (n=1,615 over 54.5 hours, 2026-08-02, provider-reported cost on every request). That is not a coding customer's day-one figure. Coding work escalates to a frontier model far more often than tool-calling does in our measurements, and escalation is most of what we pay for, so a coding-heavy workload runs closer. Send us a workload sample and we will price against your mix rather than ours.
Same work, same quality, measured against sending every call to a frontier model.
| Per 1,000 coding tasks | All-frontier | Tīrtha |
|---|---|---|
| Accuracy (HumanEval+) | 93.3% | 92.1% parity |
| Everyday work | full frontier price, every call | a fraction, routed to a lighter model |
| Your keys, your environment | ✗ | ✓ BYOK · run the light tier in-house |
| Quality ↔ cost balance | ✗ | ✓ tunable per workload |
Accuracy measured on HumanEval+ (164 problems). We state no cost multiple. The earlier "8× blended" figure is withdrawn: it was read off a meter that recorded only frontier spend, so lightweight-tier requests logged at zero and our real cost was higher than we said. Blended cost is $0.000149 per request on 1,615 of our own requests (54.5 hours to 2026-08-02, provider-reported cost on every request), not a customer workload and not coding-dominated. Not a forecast of your bill. Directional · pre-beta.
The saving above is routing, and it applies from your first request. Cache reuse is a second saving that sits on top of it. Here is exactly where that stands.
The routing saving applies from the first request, with no ramp-up. The size of it depends on how often your work escalates to a frontier model, which depends on your workload, so we measure it against your mix instead of quoting ours at you.
Spend, savings, routing mix and where each call went. Live, with nothing hidden.
Bring your own provider keys; requests are routed, not retained. The light tier can run inside your own environment, so sensitive code never has to leave.
Send us your coding volume and accuracy needs and we'll map the routing saving against your actual spend, measured, not hand-waved. Pre-beta partners get early pricing.