How we measure savings
Every competitor quotes an unfalsifiable “save up to 85%”. We would rather show the working, including the parts that are not flattering. If a number on this site cannot be traced back to something on this page, treat it as marketing and tell us.
What we do not yet claim
Progressive Labs is new and has not yet accumulated enough production traffic to publish a measured, audited savings figure. So we do not publish one. The per-request savings shown in the app are estimates, clearly labelled, calculated against the baseline model you chose. We will publish a measured median — with a confidence interval and the sampling method — once there is real traffic behind it, and not before.
1. What “quality” means here
Quality is not a single number, which is the whole reason routing works. We hold published benchmark scores per model and map them onto task types. A model’s quality for a task is expressed as a percentage of the best model available for that task — so 92% on agentic coding means 92% of the leader on agentic coding, not 92% of some global average.
Pass-rate benchmarks (SWE-bench, MMLU-Pro) normalise by ratio. Elo benchmarks normalise through win probability against the leader, which is the statistically meaningful operation on an Elo scale — averaging Elo points directly is not.
We cross-checked the two methods against each other on the coding class, using LMArena Elo and SWE-bench Verified — two entirely independent methodologies, one built on millions of human votes and one on an automated pass/fail harness. They agree within 0.4–2.6 percentage points on the models where both exist. That agreement is why we are willing to put a percentage on screen at all.
2. Where the numbers come from
Every benchmark we use, with its source and the date we last took it. Scores move; anything here older than a couple of months should be treated with suspicion, and you are welcome to hold us to that.
- SWE-bench Verifiedsourceas of 2026-08
- MMLU-Prosourceas of 2026-08
- GPQA Diamondsourceas of 2026-08
- Terminal-Bench 2.1sourceas of 2026-08
- τ²-bench Telecomsourceas of 2026-08
- LMArena Elo — codingsourceas of 2026-08-01
- LMArena Elo — mathssourceas of 2026-08-01
- LMArena Elo — creative writingsourceas of 2026-08-01
- LMArena Elo — overallsourceas of 2026-08-01
Not all sources are equal, and we weight them accordingly. GPQA’s public leaderboard is almost entirely vendor self-reported, so it carries roughly a third of the weight of an independently verified benchmark and shrinks hard toward the class average. τ²-bench is saturated — the top three models sit within 0.3 of a point — so it discriminates poorly and is used as a tie-break rather than a ranking.
A model with thin benchmark coverage is treated as uncertain, not as bad. Its estimate carries a wider band, and the quality floor is tested against the bottom of that band. The practical effect is that an unproven model has to be clearly good to be selected, which is the conservative direction.
3. How the saving is estimated
The baseline is the model you declared as your default, fixed before the request ran. Not the most expensive model in the catalogue — that would be marketing fiction, since nobody was going to run everything on it. Changing your default applies from that point forward and never rewrites history.
The estimate is honest about one specific difficulty. You cannot work out what the baseline would have cost by multiplying the observed token counts by the baseline’s rates, because a different model writes a different amount — reasoning models routinely several times more. Doing that arithmetic would systematically flatter us on verbose baselines. So we price the baseline using its own measured verbosity for that task type, and we label the result an estimate everywhere it appears: in the app, in the API response, and in the CSV export.
When routing costs more than your baseline would have — which happens, on cache misses and on tasks where the floor forces an expensive model — we show that too, as a negative. A savings ledger that can only go up is not a ledger.
4. Prices and the exchange rate
Every provider quotes in US dollars. We are a UK business and bill in pounds, so the conversion happens once, when the catalogue is priced, and everything downstream — balance, ledger, invoices — is sterling. That keeps the accounting exact and means a currency move mid-request can never change what you are charged.
The catalogue was last priced on 2026-08-07 against European Central Bank reference rates. Each currency carries its own buffer against adverse movement, and the buffer is a separate, visible number rather than something folded into the margin — so neither can quietly hide the other.
| Currency | 1 USD = | Buffer | Priced at |
|---|---|---|---|
| £ GBPBritish pound | 0.742730 | 3% | 0.765012 |
| $ USDUS dollar | 1.000000 | — | 1.000000 |
| € EUREuro | 0.865720 | 3% | 0.891692 |
| CA$ CADCanadian dollar | 1.402500 | 3% | 1.444575 |
| A$ AUDAustralian dollar | 1.419800 | 3% | 1.462394 |
The dollar row has no buffer, and that is not an oversight. Every model provider bills us in US dollars, so a dollar-denominated sale carries no currency risk at all — there is nothing to hedge, and charging for a hedge we do not need would be a quiet markup.
We price from providers’ published list prices, not from promotional rates. Several models are discounted right now; pricing off a promotion would mean our rates silently invert the moment it ends.
5. What we charge
A percentage over what the model costs us, banded by how expensive the model is: a slim markup on frontier models, which people genuinely price-shop, and a larger percentage on very cheap models, where the absolute amounts are fractions of a penny.
This is worth being direct about, because the alternative has a nasty incentive built in. Under a single flat markup, every time the router successfully moves you to a cheaper model our revenue on that request collapses — which would give us a reason to route badly. Banding removes that conflict. You still keep the overwhelming majority of the saving: a model fifty times cheaper is still tens of times cheaper after our markup.
One exception. When an identical request is served from our own cache, the upstream cost is genuinely zero and a percentage of zero is meaningless. We charge a flat tenth of what the live request would have cost, and label it as a cached response.
6. Known limitations
- Classification is heuristic today. It is fast and free and gets the certain cases exactly right, but it will mislabel some ambiguous prompts. When it is unsure it says so, and the quality floor rises in response.
- Benchmarks are proxies. A model that leads SWE-bench is not guaranteed to be better on your codebase. Compare Auto against your current model on your own work before trusting it with anything that matters.
- We have not yet measured savings by shadow-execution — running a sample of requests on your baseline as well and comparing real costs. That is the only method that fully survives “show me your evidence”, and it is what we will use before publishing any headline number.
- Published quality data lags new releases. A model released this week may have no independent benchmark coverage for weeks, which means the router treats it cautiously even if it is excellent.
Questions about any of this, or think a number is wrong? Tell us at our contact page. Corrections get made and dated.