INTEGRATION GUIDE

Python, Node.js & HTTP

Prepaid text requests

Use ZRouter from OpenAI-compatible SDKs, curl, streaming clients, and BYOK requests.

On this page

Prepare your account

Create an account, fund gateway requests with credit, and create a dedicated account API key. The hosted service is https://zrouter.si. Set these variables in your application environment, replacing the key and model placeholders:

export ZROUTER_URL="https://zrouter.si"
export ZROUTER_API_KEY="YOUR_ZROUTER_ACCOUNT_KEY"
export ZROUTER_MODEL="PROVIDER_INSTANCE/MODEL_ID"

For a separate self-hosted deployment, set ZROUTER_URL to its public URL, including any configured path prefix. Copy the model ID from the dashboard or query the inventory without generating tokens:

curl "$ZROUTER_URL/v1/models" \
  -H "Authorization: Bearer $ZROUTER_API_KEY"

Keep these credentials on the server. Browser applications should call your own backend instead of embedding an account secret in public JavaScript.

Python

Install openai in your project's environment. The client's base URL includes /v1 once:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ZROUTER_API_KEY"],
    base_url=os.environ["ZROUTER_URL"].rstrip("/") + "/v1",
    max_retries=0,
)
response = client.chat.completions.create(
    model=os.environ["ZROUTER_MODEL"],
    messages=[{"role": "user", "content": "Say hello"}],
    max_tokens=256,
)
print(response.choices[0].message.content)

Custom URLs, request options, and retry controls are documented by the official Python SDK.

Node.js / TypeScript

Install openai in your project and run this in a server-side ES module:

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.ZROUTER_API_KEY,
  baseURL: process.env.ZROUTER_URL.replace(/\/$/, "") + "/v1",
  maxRetries: 0,
});
const response = await client.chat.completions.create({
  model: process.env.ZROUTER_MODEL,
  messages: [{ role: "user", content: "Say hello" }],
  max_tokens: 256,
});
console.log(response.choices[0].message.content);

See the official JavaScript SDK. These examples disable automatic retries so you can decide how to handle potentially billable retries.

BYOK

Save a supported provider key in your account. Use a provider-qualified model belonging to that connection. Add the header on the request:

response = client.chat.completions.create(
    model=os.environ["ZROUTER_MODEL"],
    messages=[{"role": "user", "content": "Say hello"}],
    max_tokens=256,
    extra_headers={"X-ZRouter-BYOK": "true"},
)
const response = await client.chat.completions.create({
  model: process.env.ZROUTER_MODEL,
  messages: [{ role: "user", content: "Say hello" }],
  max_tokens: 256,
}, { headers: { "X-ZRouter-BYOK": "true" } });

The gateway account key remains required. The saved provider key stays in ZRouter's encrypted storage. Provider charges and any ZRouter routing fees are separate.

Streaming over HTTP

curl -N "$ZROUTER_URL/v1/chat/completions" \
  -H "Authorization: Bearer $ZROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"PROVIDER_INSTANCE/MODEL_ID","messages":[{"role":"user","content":"Say hello"}],"max_tokens":256,"stream":true}'

SDK clients can also use stream=True in Python or stream: true in JavaScript and consume the returned stream. If the final usage is unavailable after a disconnect, reserved credit can remain held for reconciliation.

Responses and embeddings

Use client.responses.create with plain text input, max_output_tokens, and store=false in Python (store: false in JavaScript). Do not pass tools, previous_response_id, a conversation ID, or background work to prepaid accounts. Send the text history explicitly for a subsequent request.

Text embeddings use /v1/embeddings with an enabled embedding model and text input. SDK methods are client.embeddings.create(model=..., input=...) in Python and client.embeddings.create({model, input}) in JavaScript. Embedding availability depends on configured providers; the model inventory may contain no embedding models.

Confirm usage

Text generation and embeddings consume tokens. Check usage, available logs, and actual ledger charges after a request. Handle 401, 402, 400, and provider errors explicitly. See shared troubleshooting.

Vendor configuration reviewed September 30, 2026. Model and client capabilities vary. Check the linked official sources when upgrading.