One API. The right model every time.
Model quality is not one number. A model that leads on agentic coding can sit mid-table on maths; one costing a fiftieth as much can be statistically level on the task actually in front of you. Progressive Labs works out which kind of work each request is, then picks the cheapest model that still clears the quality bar you set.
How a request is decided
Work out what kind of task this is
Some of it is certain rather than guessed: an attached image means vision; declared tools mean tool use; a 300,000-token prompt means long context. The rest comes from the shape and wording of the request. It costs nothing and takes under a millisecond, because anything that adds latency to every request has to justify itself against a saving measured in thousandths of a penny.
Throw away everything that cannot do the job
Context window too small, no tool support, no vision, currently rate-limited or erroring. These are hard facts, not preferences, and they are applied before price is considered at all.
Apply your quality floor
Each remaining model has a published benchmark score for this task type, normalised against the best model available for it. Anything below your floor is removed. We test against the lower end of the confidence interval, so a model never sneaks onto the shortlist on the strength of one thin, unverified number.
Only now, look at price
Among the survivors — and only among the survivors — the dial decides how hard to push for value. It cannot reach below the floor, because everything below the floor was already gone.
Tell you what happened
Which model ran, what it scored, how many alternatives cleared the bar, and what your baseline would have cost. If routing cost you more on a request than pinning would have, we show that too.
The fourteen task types
A task type earns its place only if model rankings genuinely reorder on it. Below: the best model for each, the cheapest that still clears the default floor, and the gap between them. Generated from the live catalogue of 24 models.
| Task type | Default floor | Best model | Cheapest above the floor | Price gap |
|---|---|---|---|---|
| Simple chat | 70% | Claude Fable 5 | DeepSeek V4 Flash75% | −99% |
| General chat | 82% | Claude Fable 5 | DeepSeek V4 Pro89% | −97% |
| Summarisation | 85% | Claude Fable 5 | DeepSeek V4 Pro91% | −97% |
| Extraction & classification | 85% | Claude Fable 5 | DeepSeek V4 Flash88% | −99% |
| Translation | 85% | Claude Fable 5 | DeepSeek V4 Pro89% | −97% |
| Code generation | 90% | Claude Fable 5 | Claude Sonnet 591% | −70% |
| Code review & debugging | 92% | Claude Fable 5 | Claude Opus 593% | −50% |
| Agentic coding | 96% | Claude Fable 5 | Claude Fable 5100% | same model |
| Maths & logic | 90% | Claude Opus 5 | Qwen3.7 Plus99% | −94% |
| Science & research | 92% | Claude Fable 5 | DeepSeek V4 Flash97% | −99% |
| Long-context work | 90% | Claude Fable 5 | DeepSeek V4 Pro91% | −97% |
| Tool use & agents | 94% | Claude Fable 5 | Claude Fable 5100% | same model |
| Creative writing | 85% | Claude Fable 5 | Qwen3.8 Max96% | −84% |
| Vision | 88% | Claude Fable 5 | DeepSeek V4 Pro90% | −97% |
“Price gap” compares blended per-token rates, weighted three-to-one toward input because that is how real traffic runs. It is the headroom available on that task type, not a promise about your particular workload — what you actually save depends on what you actually send.
What it will never do
See it decide on your own work
Put Auto against the model you use today. If we do not beat it on your real prompts, you have lost the price of a few tokens finding out.
Compare against your model