Skip to content

Documentation

One OpenAI-compatible endpoint for every model in the catalogue. Change the base URL and the model string; the rest of your code stays as it is.

Quickstart

Create a key from your dashboard, then point any OpenAI-compatible client at https://progressivelabs.uk/api/v1.

curl
curl https://progressivelabs.uk/api/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $PROGRESSIVE_LABS_API_KEY" \
  -d '{
    "model": "auto",
    "messages": [{"role": "user", "content": "Explain KV caching in one paragraph."}],
    "stream": true
  }'

Passing auto as the model hands the choice to the router: it works out what kind of task the request is and picks the cheapest model that still clears your quality floor. Pass a specific slug instead and that exact model runs, every time.

Controlling the router

Three optional parameters. All of them are ignored when you pin a specific model.

json
{
  "model": "auto",

  // 0 = best quality at any price, 10 = cheapest above the floor. Default 6.
  "costQualityDial": 6,

  // Hard minimum quality, as a fraction of the best model for the detected
  // task. Omit to use our per-task defaults (0.70 for trivial chat through
  // 0.96 for agentic coding). This is a constraint the dial cannot cross.
  "qualityFloor": 0.9,

  // The model you would otherwise have used. Every saving we report is
  // measured against this, and it is fixed before the request runs.
  "baselineModel": "anthropic/claude-fable-5",

  // Keeps a conversation on one model so its prompt cache stays warm.
  "sessionId": "conversation-abc123"
}

Every streamed response carries a route event before the first token, telling you which model was chosen, what task type was detected, what it scored, and how many alternatives cleared the bar. The closing done event carries the cost, the estimated baseline cost and the difference.

Python

Works with the official OpenAI SDK — only base_url changes.

python
from openai import OpenAI

client = OpenAI(
    base_url="https://progressivelabs.uk/api/v1",
    api_key=os.environ["PROGRESSIVE_LABS_API_KEY"],
)

stream = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Summarise this changelog."}],
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

TypeScript

typescript
import OpenAI from 'openai'

const client = new OpenAI({
  baseURL: 'https://progressivelabs.uk/api/v1',
  apiKey: process.env.PROGRESSIVE_LABS_API_KEY,
})

const stream = await client.chat.completions.create({
  model: 'auto',
  messages: [{ role: 'user', content: 'Write a regex for UK postcodes.' }],
  stream: true,
})

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? '')
}

Model identifiers

Models use provider/model notation. Switching providers is a one-string change.

text
anthropic/claude-fable-5
anthropic/claude-opus-5
anthropic/claude-sonnet-5
alibaba/qwen3-7-max
deepseek/deepseek-v3-2
zai/glm-5-2

Billing and limits

  • Every request is priced from actual token usage and debited from your balance.
  • A request is rejected with 402 if your balance cannot cover its maximum possible cost.
  • Usage headers on every response report the tokens consumed and the charge applied.
  • Rate limits are per key and per account; see your dashboard for current values.

Data handling

Prompts and completions are not retained. We store a salted hash of each prompt for abuse detection, plus token counts and cost. See the privacy policy for the full detail.

Something missing? Tell us or email hello@progressivelabs.uk.