Skip to content
← Blog
June 25, 2026

Why Kumo's rate lands 30–50% below list — and how to check it

pricingtrust
pricing

When a gateway says "30–50% less", the reasonable question is: less than what, and is it really the same model? Here's the honest version, without the marketing gloss.

It's the same models, guaranteed

Same frontier models. A request for GPT-5.4, Claude Sonnet 4.6 or Gemini 2.5 Pro reaches that exact model — the provider's response comes back verbatim, and every response reports which model served it. The lower rate never comes from a quiet substitution.

An effective rate, not a flat discount

Kumo quotes an effective rate — what you actually pay per token, all-in, after every optimization Kumo applies automatically on your behalf. It moves with your usage pattern and volume, which is why you see the exact number for each model before you spend, instead of one flat headline you'd have to take on faith.

The honest range

The 30–50% figure is real, but it isn't identical for every account — your workload and volume move it within that band. What's fixed: you always see the exact rate for the exact model before you spend a token, and that rate tracks the underlying provider's list price, so when a lab cuts its price, yours drops too.

How to verify it yourself

Pin a model and it's guaranteed — you get that model, not a substitute. Every response reports which model served it, so you can check that field yourself rather than take our word for it.

Start building on Kumo today

One base URL, one balance, every model — at an effective rate you can see before you spend a token.