When a gateway says "30–50% less", the reasonable question is: less than what, and is it really the same model? Here's the honest version, without the marketing gloss.
It's the same models, guaranteed
Same frontier models. A request for GPT-5.4, Claude Sonnet 4.6 or Gemini 2.5 Pro reaches that exact model — the provider's response comes back verbatim, and every response reports which model served it. The lower rate never comes from a quiet substitution.
An effective rate, not a flat discount
Kumo quotes an effective rate — what you actually pay per token, all-in, after every optimization Kumo applies automatically on your behalf. It moves with your usage pattern and volume, which is why you see the exact number for each model before you spend, instead of one flat headline you'd have to take on faith.
The honest range
The 30–50% figure is real, but it isn't identical for every account — your workload and volume move it within that band. What's fixed: you always see the exact rate for the exact model before you spend a token, and that rate tracks the underlying provider's list price, so when a lab cuts its price, yours drops too.
How to verify it yourself
Pin a model and it's guaranteed — you get that model, not a substitute. Every response reports which model served it, so you can check that field yourself rather than take our word for it.