One API key. Top models. Zero rewrites.
TokenAPI is an OpenAI-compatible gateway. Call stable model IDs with the SDK you already use; we route every request to a strong backend, fail over automatically and bill you per token.
Your code talks to one endpoint. We handle the rest.
A model ID is a stable product, not a hard-wired vendor model. Behind it sits a primary route and ordered fallbacks. Upgrading or re-routing a model happens on our side; your requests stay exactly the same.
- 1
Your app
model: "tokenapi-pro"to/v1/chat/completions - 2
TokenAPI gateway
Authenticates the key, applies limits, reserves the maximum cost.
- 3
Routing
Primary backend first; on failure, the next route takes over before any token is sent.
- 4
Response
Streamed back in the format you called, then settled to the exact token usage.
Built for teams shipping AI to production.
Works with the SDKs you already use
Point the official OpenAI or Anthropic SDK at our base URL. No new client library, no code rewrite.
Model names that never break your code
You call a stable model ID such as tokenapi-pro. We keep it served by a strong backend, so your integration never has to change.
Automatic failover
If a backend is slow, rate-limited or down, the request is retried on the next route before you receive a single token.
Server-sent events everywhere
Token-by-token streaming on all three API formats, including tool calls and usage reporting.
Exact, per-request accounting
Every request is metered in micro-dollars. A request reserves its maximum cost up front and settles to the real usage, so a balance is never overdrawn.
Limits per API key
Per-key rate limits, spend caps, expiry dates and model allow-lists keep each project and teammate in bounds.
Pay as you go. No subscription.
Per-token pricing
Each model has its own input and output price per million tokens. Top up a balance and spend it on any model.
- Prices listed per model
- Charged on actual tokens
- Unused reservations released instantly
Custom
Higher rate limits, invoicing and volume pricing for production workloads.
- Custom rate limits
- Multiple keys with spend caps
- Priority support
Common questions.
Which models can I use?
The live list, with capabilities, context length and prices, is on the Models page. Each TokenAPI model is served by a leading upstream model provider; we may change the backend behind a model ID to keep quality and uptime high. Ask any model what it runs on and it will tell you.
Do I have to change my code?
Only the base URL and the API key. Requests and responses follow the OpenAI Chat Completions and Responses formats and the Anthropic Messages format.
What happens when a provider has an outage?
Models can have several backend routes. If the primary fails before it starts answering, the next route is tried automatically and you get a normal response.
How am I billed?
Prepaid balance, per-token prices listed per model. Before a request runs we reserve its maximum possible cost, then charge the actual tokens and release the rest immediately.
Do you store my prompts?
No. We log request metadata needed for billing and support (model, token counts, cost, latency, status), not prompt or completion content. Requests are processed by the upstream provider serving the model.
How do I get an API key?
We are onboarding customers manually during early access. Email hello@tokenapi.biz and we will set up your account, balance and keys.
Get your API key.
Early access is open. Tell us what you are building and we will set up your account, balance and keys.