API reference

Every LokaRouter gateway endpoint — OpenAI-compatible request and response shapes.

Authentication

All gateway endpoints accept your API key as a Bearer token in the Authorization header. Create keys in the dashboard; keys are scoped to a workspace and can carry guardrails.

Authorization: Bearer sk-lkr-your-key

Endpoints

  • POST/v1/chat/completions— Create a chat completion (streaming supported via stream: true)
  • POST/v1/completions— Create a legacy text completion (streaming supported)
  • POST/v1/embeddings— Create embedding vectors for text input
  • GET/v1/key— Inspect this key — remaining budget and rate limit (free, no tokens used)
  • GET/v1/credits— Account credits — balance and this month's usage (free, no tokens used)
  • GET/v1/generation?id=req_…— Look up one request's cost, tokens, and model by x-request-id
  • GET/v1/models— List models available through the gateway
  • GET/v1/models/{id}— Model detail — context window, pricing, tags
  • GET/v1/files— List your uploaded files (Beta)
  • POST/v1/files— Upload a file (multipart/form-data or JSON base64, 5 MB cap)
  • GET/v1/files/{id}— Retrieve file metadata
  • GET/v1/files/{id}/content— Download file contents
  • DELETE/v1/files/{id}— Delete a file

Chat completions

POST /v1/chat/completions with any OpenAI-compatible payload. Unknown fields (temperature, tools, response_format, …) are forwarded to the upstream provider untouched. Set "stream": true for server-sent events.

curl https://api.lokarouter.id/v1/chat/completions \  -H "Authorization: Bearer sk-lkr-your-key" \  -H "Content-Type: application/json" \  -d '{    "model": "openai/gpt-5-mini",    "messages": [{ "role": "user", "content": "Hello" }],    "stream": false  }'

Response

{  "id": "lkr_…",  "object": "chat.completion",  "model": "openai/gpt-5-mini",  "choices": [{    "index": 0,    "message": { "role": "assistant", "content": "Hi!" },    "finish_reason": "stop"  }],  "usage": { "prompt_tokens": 9, "completion_tokens": 2, "total_tokens": 11 }}

Every response carries x-request-id (per-request lookup) and x-lokarouter-latency-ms (end-to-end gateway latency).

Auto-routing

Set "model": "lokarouter/auto" and the gateway classifies your request heuristically (code, translation, long context, reasoning, latency-sensitive) and routes it to the best catalog model. Costs nothing extra and adds no latency.

{  "model": "lokarouter/auto",  "messages": [{ "role": "user", "content": "Refactor this function…" }]}

The chosen model is what gets recorded in usage logs. Allow-list lokarouter/auto in a key's guardrails to permit every concrete model through auto.

Files (Beta)

Store small files (up to 5 MB) for reuse — e.g. fine-tune or batch inputs. Purposes: batch, fine-tune, vision, assistants, user_data.

curl https://api.lokarouter.id/v1/files \  -H "Authorization: Bearer sk-lkr-your-key" \  -F "file=@training.jsonl" \  -F "purpose=fine-tune"

Listing files

curl https://api.lokarouter.id/v1/files \  -H "Authorization: Bearer sk-lkr-your-key"

Rate limiting

Keys allow 60 requests per minute by default. When the limit is hit the gateway returns 429 with a retry-after header plus x-ratelimit-limit and x-ratelimit-remaining so SDKs can back off correctly.

Errors

Errors use the OpenAI shape — an error object with message, type, and code. Common codes:

{  "error": {    "message": "Key budget exhausted: $5.00 spent of $5.00 budget.",    "type": "insufficient_quota",    "code": "budget_exhausted"  }}
HTTPCodeMeaning
400invalid_json / missing_required_parameterMalformed request — bad JSON or missing model/messages.
401invalid_api_keyMissing or invalid API key.
402budget_exhausted / workspace_budget_exhaustedKey or workspace budget exhausted.
403model_not_allowed / tool_not_allowedModel or tool blocked by the key's guardrails.
404model_not_found / file_not_foundUnknown model id or file id.
429rate_limit_exceededRate limit exceeded — retry after the hinted delay.
502provider_failed / upstream_unavailableAll upstream providers failed or none are configured.

← Back to Docs