TokenAPIAPI docs
API reference

Integrate in two lines: base URL and API key.

TokenAPI speaks the OpenAI Chat Completions and Responses formats and the Anthropic Messages format. Use the official SDKs unchanged and pick a model ID from the model list.

Quickstart

Base URL

https://tokenapi.biz/v1

All endpoints:

GET/v1/models

List models

Public model IDs with capabilities, context length and prices.

GET/v1/models/{id}

Retrieve a model

A single model object.

POST/v1/chat/completions

Chat Completions

OpenAI Chat Completions format, streaming and non-streaming.

POST/v1/responses

Responses

OpenAI Responses format (stateless), streaming and non-streaming.

POST/v1/messages

Messages

Anthropic Messages format, for the Anthropic SDK.

curl
curl https://tokenapi.biz/v1/chat/completions \
  -H "Authorization: Bearer $TOKENAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tokenapi-pro",
    "messages": [{ "role": "user", "content": "Hello!" }]
  }'
Node.js · openai
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://tokenapi.biz/v1",
  apiKey: process.env.TOKENAPI_KEY,
});

const res = await client.chat.completions.create({
  model: "tokenapi-pro",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(res.choices[0].message.content);
Python · openai
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokenapi.biz/v1", api_key=os.environ["TOKENAPI_KEY"])

res = client.chat.completions.create(
    model="tokenapi-pro",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(res.choices[0].message.content)
Authentication

Send your key as a Bearer token.

Authorization: Bearer sk-tk-…

The Anthropic-style x-api-key header is accepted too. Keys start with sk-tk-, are shown once at creation and can be limited per key (requests per minute, spend cap, allowed models, expiry). Keep keys on the server; never ship them in browser or mobile code. To get a key, email hello@tokenapi.biz.

Models

Stable model IDs.

Use the id from GET /v1/models as the model parameter. A model ID is a stable product: we may upgrade or re-route the backend that serves it to keep quality and availability high, without any change on your side. Responses always report the model ID you called.

GET /v1/models
{
  "object": "list",
  "data": [
    {
      "id": "tokenapi-pro",
      "object": "model",
      "owned_by": "tokenapi",
      "display_name": "TokenAPI Pro",
      "capabilities": ["streaming", "tools", "json_output"],
      "context_length": 131072,
      "pricing": {
        "currency": "USD",
        "prompt": "0.0000015",
        "completion": "0.000006",
        "input_per_million": 1.5,
        "output_per_million": 6
      }
    }
  ]
}

pricing.prompt / pricing.completion are USD per token (strings); input_per_million / output_per_million are USD per million tokens.

POST /v1/chat/completions

Chat Completions

Request and response follow the OpenAI format: messages with text and image parts, tools / tool_choice, response_format, temperature, top_p, stop, max_tokens or max_completion_tokens. Each response carries an x-request-id header; include it when contacting support.

If you omit max_tokens, the model's maximum output length is applied.

Streaming

Server-sent events

Set stream: true. Chunks arrive as data: lines and the stream ends with data: [DONE]. Add stream_options.include_usage to receive token usage in a final chunk.

Node.js
const stream = await client.chat.completions.create({
  model: "tokenapi-pro",
  messages: [{ role: "user", content: "Write a haiku." }],
  stream: true,
  stream_options: { include_usage: true }, // final chunk carries usage
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
POST /v1/responses

Responses API

Supports string or array input, instructions, function tools, text.format (JSON schema) and the full streaming event sequence (response.output_text.delta, response.completed, …).

Node.js
const res = await client.responses.create({
  model: "tokenapi-pro",
  instructions: "Answer in one sentence.",
  input: "What is a token?",
});
console.log(res.output_text);
POST /v1/messages

Anthropic Messages

For code written against the Anthropic SDK. Supports system, text and image blocks, tools with tool_use / tool_result, and streaming events. Errors use the Anthropic error shape.

Node.js · @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";

const anthropic = new Anthropic({
  baseURL: "https://tokenapi.biz",   // SDK appends /v1/messages
  apiKey: process.env.TOKENAPI_KEY,
});

const msg = await anthropic.messages.create({
  model: "tokenapi-pro",
  max_tokens: 1024,
  messages: [{ role: "user", content: "Hello!" }],
});
Errors

OpenAI-style error objects

Example
{
  "error": {
    "message": "The model 'nope' does not exist or is disabled.",
    "type": "invalid_request_error",
    "code": "model_not_found",
    "param": null
  }
}
400invalid_request_error

Malformed JSON, missing model/messages, or an unsupported parameter.

401invalid_api_key

Missing, wrong, disabled or expired API key.

402insufficient_quota

Balance cannot cover this request's maximum cost, or the key's spend cap is reached. Top up or lower max_tokens.

403model_not_allowed

This key is restricted to other models.

404model_not_found

Unknown or disabled model ID. List valid IDs with GET /v1/models.

429rate_limit_exceeded

Per-key requests-per-minute limit exceeded. Retry after the Retry-After header.

502upstream_unavailable

Every backend route for this model failed. Safe to retry.

Upstream failures are retried on another route automatically before you see an error. A 502 means all routes failed and you were not charged.

Billing & limits

Prepaid, per token, never overdrawn.

  • Prices are per model, in USD per million input and output tokens (see Models).
  • Before a request runs, its maximum possible cost (prompt plus max_tokens) is reserved from your balance. If the balance cannot cover it, you get a 402 and nothing is sent upstream.
  • When the request finishes, you are charged the actual tokens and the rest of the reservation is released immediately.
  • Requests that fail before the model starts answering (any 4xx/5xx error response) are not charged. If a stream is cancelled or breaks midway, only the tokens already generated are charged.
  • Rate limits are per API key, in requests per minute; exceeding them returns 429 with Retry-After.
Compatibility notes

Not supported yet

  • /v1/responses: previous_response_id (send the full conversation instead), built-in tools such as web_search, and file_id inputs.
  • /v1/messages: extended thinking and cache_control are ignored.
  • n > 1 on Chat Completions.
Questions? Contact us