Skip to content
24 models · 10 providers · priced in GBP

The right model for every task, at the lowest price that keeps the quality.

Some models are far better at code than at maths. Some are a fiftieth of the price and just as good at the job in front of them. We work out which is which, per request, and send your work to the one that wins — while you keep a hard floor on quality that we are not allowed to cross.

All systems operational
14Task types detected
24Models, one API
<1msRouting overhead
NoSubscription required

The same six requests, routed

Every row below is generated by running the actual router when this page is built — not written by a marketer. The baseline is Claude Fable 5, the sort of model you might reasonably default to for everything.

RequestDetected asRouted toQualityCostvs baseline
Acknowledging a messageSimple chatDeepSeek V4 ProDeepSeek85%£0.000007−97%
Summarising a long documentSummarisationQwen3.7 PlusAlibaba90%£0.001176−97%
Writing a new functionCode generationClaude Opus 5Anthropic94%£0.001492−53%
Debugging a stack traceCode review & debuggingClaude Opus 5Anthropic93%£0.00199−53%
A hard maths problemMaths & logicClaude Opus 5Anthropic100%£0.000498−53%
Drafting a short storyCreative writingQwen3.8 MaxAlibaba96%£0.000135−87%

Quality is that model’s published benchmark score on the detected task type, as a percentage of the best model available for that task. Costs are estimates for a 1,024-token reply — a different model writes a different number of words, so the exact figure moves. The median saving across these six is 87%. We publish the whole method, including where the benchmark numbers come from and what we cannot prove yet.

Cheaper is only worth it
if the answer still lands

A router that quietly downgrades your output to save a penny is worse than no router. So the price dial in Progressive Labs is not allowed to touch quality at all.

How the floor works
01

The floor is a constraint, not a preference

You set a minimum — "never below 92% of the best model for this task". Models that fail it are removed from consideration before price is even looked at. The dial only ever chooses between survivors.

02

When we are unsure, we get more expensive

If the classifier is not confident about what kind of work this is, the floor rises automatically. An uncertain router should be a cautious one — misrouting an agentic coding task costs you a failed task, which dwarfs the tokens it saved.

03

You can always pin a model

Choose a specific model and that is exactly what runs, every time, with no second-guessing. We still show you afterwards what routing would have done, so the choice stays yours and stays informed.

Then we cut the tokens
themselves

Picking a cheaper model is only half of it. The other half is not paying for the same tokens twice.

Repeat answers cost a tenth

Ask the identical question twice at temperature zero and the second one is served from cache. It costs the upstream provider nothing, so we charge a flat 10% instead of full price rather than pretending we did the work.

Warm prompt caches, managed for you

Long system prompts and conversation history can be cached upstream at up to 99% off. We track which prefixes actually repeat and only pay the cache-write premium when the maths says it will earn its money back.

Routing that knows about the cache

Switching model throws away a warm cache. The router prices that in, so it will not move you to save 15% on paper while quietly losing a 90% discount.

One conversation, one model

Within a session we keep you on the model we picked, so the prefix stays warm. Consistency is usually worth more than chasing the cheapest option on every turn.

What you get

One account, one balance, in pounds.

Smart routing

Send your request to auto and we choose the model. You see which one ran, why, what it scored, and what it saved.

Compare

Two models, one prompt, one screen — or put Auto against your current default and see whether we actually beat it on your real work.

API

One OpenAI-compatible endpoint. Change the base URL, pass auto as the model, and keep the rest of your code exactly as it is.

Usage & billing

Every run itemised: which model, which task type, what it scored, what it cost, and what your baseline would have cost. Exportable as CSV.

Fair by construction

The things we decided up front so you don't have to take our word for them later.

Prepaid credit, never a surprise bill

You top up; we draw down. There is no invoice at the end of the month and no way to spend money you have not already put in.

Every charge is itemised

Each run writes a ledger entry with the running balance, so any charge reconciles back to the exact request that caused it.

Prompts are not stored

We keep a hash for abuse detection and nothing else. Your prompts and completions are not retained, not logged, and never used for training.

We show our working

Benchmark sources, dates, the normalisation formula and the limits of what we can prove are all published. A saving you cannot audit is a saving you should not believe.

Models we route between

Retail prices per million tokens in GBP. The full catalogue lists every model we carry.

View all 24 models

Claude Fable 5

AnthropicAnthropic's most capable model. Top of SWE-bench Verified and top of the creative-writing arena — the one to reach for when being right matters more than what it costs.Input /M£9.18Output /M£45.90

Claude Opus 5

AnthropicComplex agentic coding and long-horizon work at half the price of Fable 5. Leads Terminal-Bench 2.1 and the maths arena.Input /M£4.59Output /M£22.95

Claude Sonnet 5

AnthropicNear-Opus coding quality at Sonnet money — 0.852 on SWE-bench Verified for a fifth of Fable 5. One of the best value-per-point models in the catalogue.Input /M£2.87Output /M£14.34

Qwen3.8 Max

AlibabaAlibaba's newest flagship — 2.4T-parameter MoE, natively multimodal, third in the creative-writing arena. Flat pricing across the whole 1M window, with no long-context surcharge.Input /M£1.91Output /M£5.74

GPT-5.6 Sol

OpenAIOpenAI's flagship. Leads Terminal-Bench 2 and tops the GPQA science leaderboard.Input /M£4.59Output /M£27.54

Gemini 3.6 Flash

GooglePunches far above its price on maths — fourth in the arena, level with models costing several times more.Input /M£1.43Output /M£7.17

DeepSeek V4 Pro

DeepSeekFrontier-adjacent quality at a tenth of frontier prices, and by far the most aggressive prompt-cache rate on the market — cached input costs under 1% of fresh.Input /M£0.466Output /M£0.932

MiniMax M3

MiniMaxStatistically level with Gemini 3.1 Pro on SWE-bench Verified at a fraction of the price. The clearest example of why routing pays.Input /M£0.643Output /M£2.57

Stop paying frontier prices
for everyday work.

Free to sign up. No card required to look around. Add credit only when you want to run something.