ПОПУЛЯРНЫЙ
60,000,000 tokens
≈20,000 API requests on a single balance
$0.34 / 1M tokens$20.62
≈ 1 639 ₽
One API for supported GPT, Claude and Gemini models — one key, one balance, prices 30–50% lower. Every model is the original — no swaps, no downgraded versions.
No credit card · one base URL · no VPN · free to start
Models supported
Models are provided by their labs. Kumo is not affiliated with or endorsed by them.
Packages
One balance, one key. Top up and pay per token — no subscriptions, no minimums.
ПОПУЛЯРНЫЙ
≈20,000 API requests on a single balance
$0.34 / 1M tokens$20.62
≈ 1 639 ₽
Pay in rubles — by card or SBP; companies get an invoice with closing documents.
Full pricing & calculatorSide projects and experiments
API access
$2.25
≈ 179 ₽
API access
$3.38
≈ 269 ₽
API access
$5.02
≈ 399 ₽
A key inside your product — lower token COGS
API access
$6.78
≈ 539 ₽
API access
$13.83
≈ 1 099 ₽
API access
$18.10
≈ 1 439 ₽
Pooled team volume, one invoice
API access
$20.62
≈ 1 639 ₽
API access
$30.18
≈ 2 399 ₽
API access
$47.80
≈ 3 799 ₽
Models
A curated set of production-ready GPT, Claude and Gemini picks at Kumo's effective rate.
Same model, same answer — 30% lower effective rate.
Get your API keySame model, same answer — 30% lower effective rate.
Get your API keySame model, same answer — 30% lower effective rate.
Get your API keyFrom signup to first call
Sign up, top up once, grab a key — and you're calling supported GPT, Claude and Gemini text models from one balance. No sales call, no contract, no per-seat fees.
Create an account to get started. Add your team to an org whenever you're ready.
Buy credits once, then spend them on any model or provider from a single balance.
Create a key and start shipping. Fully OpenAI-compatible — change one base URL.
Generate a key to see your token budget.
Who it's for
Developers, AI agents, products, enterprise — we shape Kumo around how each one uses tokens: the same models for less, fast and without a VPN, and the difference goes straight back into your margin.
What you get
Drop-in via the OpenAI API: switch one base URL, keep the rest of your code — Cursor, Cline, Roo, Continue or your own scripts. You pick every model yourself: low-cost DeepSeek and Qwen, flagship Claude or GPT — call whatever fits the task, billed at Kumo's optimized rate. Pay only for real tokens — no seats, no subscription, no VPN.
Volume / month
↳ that's ≈ 5 features (~1K LOC) · ≈ 3 days in Claude Code
Indicative estimate at a 75% input / 25% output split. "List" is the provider's standard rate; "Your price" is Kumo's automatically optimized effective rate. Tokens per unit are conservative estimates. Exact rates are visible in your dashboard before you're charged.
Security & trust
Security and reliability are built into Kumo from day one. Here's how we protect your data and keep the service dependable.
Security & data
Prompts and completions are never used to train models. Content is processed to serve your request and nothing more.
We store usage metadata for billing — not the body of your prompts. Bring your own retention policy.
Kumo forwards prompts and completions to serve the call, then drops the content. We retain only the account, billing and usage metadata needed to operate the service.
Why Kumo
OpenAI, Anthropic and Google models behind a single API key. If a provider errors, the gateway returns a 502 and you can switch providers directly.
Real-time spend, per-key limits and detailed logs. Overspend shows up the moment it happens — no surprise invoice at month's end.
Works with Claude Code, Cline, Cursor, Kilo Code, Opencode and any OpenAI SDK — no code changes. Swap one base URL and you're done.
How people use Kumo
From a solo build on nights and weekends to agents running 24/7 and an agency's whole AI stack — here's how different teams actually run on Kumo.
Next steps
The full picture of how Kumo runs under the hood, plus a direct line for volume deals and consulting.
We reply within one business day.
Moving serious volume? Volume discounts, company invoicing and hands-on consulting — talk to us