How it works
One gateway. Every model. Yours in minutes.
Kumo sits in front of supported GPT, Claude and Gemini text models as a single OpenAI-compatible endpoint. You change one base URL; we handle routing, billing and access — and pass the optimized rate back to you.
How the service works
What happens on every call.
You call one endpoint
Point your existing OpenAI SDK at api.kumo.cloud/v1 and send the request you already wrote. No new client, no per-model SDK.
We bill it at the optimized rate
Every call is billed at Kumo's optimized effective rate for the exact model you asked for — no plan to pick, no lever to flip. Pin a model and you get that model, never a downgrade.
You get the provider's answer, verbatim
The response is the upstream provider's, unchanged. Every call reports which model served it — so you can verify it yourself, not just take our word for it.
It all bills from one balance
Tokens, requests, latency and spend stream into one dashboard. Cap spend, set alerts, top up once — and stop reconciling five vendor invoices.
What we're built on
Real accounts. Real models. No smoke.
Kumo isn't a wrapper around a wrapper. We hold direct, paid accounts with the labs themselves — that's why the catalog you see is the catalog you actually call, at Kumo's optimized rate.
Direct lab accounts
We connect straight to OpenAI, Anthropic and Google on our own committed contracts — not a chain of resellers. Your call reaches the real provider, every time.
One live catalog
The 9 active text models and prices you see are the ones we actually serve — the same list the pricing table reads from. No phantom models, no bait-and-switch.
Rates that track the market
Effective rates track the underlying provider list price — when a lab cuts a price, yours drops too.
Multi-provider access
Because we hold accounts across providers, the catalog spans multiple labs. If a provider errors, you get a 502 and can retry or switch providers directly.
Why it's cheaper — and still legit
The same models, priced transparently.
You're not getting a knockoff. It's GPT, Claude and Gemini, billed through Kumo — at a lower effective rate, with the exact price shown before you spend.
Same models, real accounts
GPT, Claude and Gemini, billed through Kumo — not a knockoff, not a downgrade.
Rate shown before you spend
See the exact per-model rate up front. No hidden line items, no surprise multipliers.
Buy in bulk, save more
Commit to a volume up front and lock a lower rate for the term.
Never a silent swap
Pin a model and you get that exact model. Every response is the provider's, verbatim.
lower effective token bill
Your exact saving depends on your usage pattern and volume — you always see the per-model rate before you spend.
Same model IDs, the provider's response verbatim — you ask for a model, you get that model, never a silent substitution. Pin it and it's guaranteed, and every call reports which model served it. So don't take our word for it — you can check.
Independent research: how to audit an API for a swapped or weaker model →Drop-in compatible
One line in. Savings out.
A real, playable terminal — type a command or click a step. The whole migration, the call, and the saved number, live.
Security & encryption
Your data stays yours.
Kumo is a billing-and-routing layer, not a data lake. We process only what's needed to serve and meter the request — nothing more.
Encrypted in transit & at rest
Every request runs over TLS 1.2+; stored metadata and secrets are encrypted at rest with AES-256. API keys are stored hashed and shown once, at creation.
Never trained on
Prompts and completions are never used to train models. They're forwarded to serve your request, then dropped.
Prompt bodies aren't logged
We keep usage metadata for billing — token counts, model, timestamp — not the content of your prompts or responses. Bring your own retention policy.
Scoped, revocable keys
Every key is scoped and revocable, shown once at creation.
FAQ
Questions, answered.
Everything you need to know about buying and spending tokens on Kumo.
Spin up your gateway today.
One base URL, one balance, your whole team behind it. Free to start.