Skip to content

Making Requests ​

Learn how to make API requests to Tokenlio, handle responses, and implement best practices for production applications.

Base URL ​

All API requests should use:

https://api.tokenlio.ai/v1

Request Format ​

Tokenlio follows the OpenAI API specification. If you're familiar with OpenAI's API, you're already familiar with Tokenlio.

Basic Request Structure ​

bash
curl https://api.tokenlio.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "gpt-4-turbo",
    "messages": [
      {"role": "user", "content": "Hello!"}
    ]
  }'

Required Headers ​

  • Content-Type: application/json - JSON request body
  • Authorization: Bearer YOUR_API_KEY - Your API key

Response Format ​

Success Response ​

json
{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1699999999,
  "model": "gpt-4-turbo",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! How can I assist you today?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 10,
    "completion_tokens": 12,
    "total_tokens": 22
  }
}

Error Response ​

json
{
  "error": {
    "message": "Insufficient balance",
    "type": "insufficient_balance",
    "code": "insufficient_balance"
  }
}

Common Parameters ​

Model Selection ​

Specify which model to use:

json
{
  "model": "gpt-4-turbo"
}

See available models for the full list.

Messages Array ​

Chat completions use a messages array:

json
{
  "messages": [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "What is AI?"},
    {"role": "assistant", "content": "AI stands for..."},
    {"role": "user", "content": "Tell me more"}
  ]
}

Roles:

  • system: Sets assistant behavior (optional, but recommended)
  • user: User messages
  • assistant: Previous assistant responses (for context)

Temperature ​

Controls randomness (0.0 to 2.0):

json
{
  "temperature": 0.7
}
  • 0.0: Deterministic, focused
  • 0.7: Balanced (default)
  • 1.0+: More creative, random

Max Tokens ​

Limit output length:

json
{
  "max_tokens": 500
}

TIP

Set max_tokens to control costs. The model stops generating when this limit is reached.

Other Parameters ​

json
{
  "top_p": 0.9,           // Alternative to temperature
  "frequency_penalty": 0, // Reduce repetition (-2.0 to 2.0)
  "presence_penalty": 0,  // Encourage new topics (-2.0 to 2.0)
  "stop": ["\n", "END"]   // Stop sequences
}

Streaming Responses ​

Get tokens as they're generated instead of waiting for completion.

Python ​

python
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.tokenlio.ai/v1"
)

stream = client.chat.completions.create(
    model="gpt-4-turbo",
    messages=[{"role": "user", "content": "Tell me a story"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

JavaScript ​

javascript
const stream = await client.chat.completions.create({
  model: 'gpt-4-turbo',
  messages: [{ role: 'user', content: 'Tell me a story' }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content || '');
}

cURL (Server-Sent Events) ​

bash
curl https://api.tokenlio.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "gpt-4-turbo",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": true
  }'

Error Handling ​

Status Codes ​

CodeMeaningCommon Causes
200SuccessRequest completed
400Bad RequestInvalid parameters
401UnauthorizedInvalid API key
402Payment RequiredInsufficient balance
403ForbiddenModel not accessible
429Rate LimitedToo many requests
500Server ErrorInternal error
503Service UnavailableTemporary outage

Python Error Handling ​

python
from openai import OpenAI, OpenAIError, RateLimitError, AuthenticationError

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.tokenlio.ai/v1"
)

try:
    response = client.chat.completions.create(
        model="gpt-4-turbo",
        messages=[{"role": "user", "content": "Hello"}]
    )
except AuthenticationError:
    print("Invalid API key")
except RateLimitError:
    print("Rate limit exceeded - slow down")
except OpenAIError as e:
    print(f"API error: {e}")

JavaScript Error Handling ​

javascript
try {
  const response = await client.chat.completions.create({
    model: 'gpt-4-turbo',
    messages: [{ role: 'user', content: 'Hello' }],
  });
} catch (error) {
  if (error.status === 401) {
    console.error('Invalid API key');
  } else if (error.status === 429) {
    console.error('Rate limit exceeded');
  } else {
    console.error('API error:', error);
  }
}

Retry Logic ​

Implement exponential backoff for transient errors:

python
import time
from openai import OpenAI, RateLimitError

def make_request_with_retry(client, max_retries=3):
    for attempt in range(max_retries):
        try:
            return client.chat.completions.create(
                model="gpt-4-turbo",
                messages=[{"role": "user", "content": "Hello"}]
            )
        except RateLimitError:
            if attempt == max_retries - 1:
                raise
            wait_time = (2 ** attempt)  # 1s, 2s, 4s
            print(f"Rate limited. Waiting {wait_time}s...")
            time.sleep(wait_time)

Request IDs ​

Every response includes a unique ID:

json
{
  "id": "chatcmpl-abc123",
  ...
}

Use this ID when contacting support - it helps us debug issues quickly.

Best Practices ​

1. Set Timeouts ​

Prevent hanging requests:

python
import httpx

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.tokenlio.ai/v1",
    timeout=30.0,  # 30 second timeout
    http_client=httpx.Client()
)

2. Use System Messages ​

Guide model behavior:

json
{
  "messages": [
    {
      "role": "system",
      "content": "You are a technical support assistant. Be concise and helpful."
    },
    {"role": "user", "content": "How do I reset my password?"}
  ]
}

3. Validate Input ​

Check user input before sending:

python
def validate_message(content: str) -> bool:
    if not content or not content.strip():
        return False
    if len(content) > 10000:  # Max length
        return False
    return True

4. Cache Responses ​

Avoid duplicate API calls:

python
from functools import lru_cache

@lru_cache(maxsize=100)
def get_completion(prompt: str) -> str:
    response = client.chat.completions.create(
        model="gpt-4-turbo",
        messages=[{"role": "user", "content": prompt}],
        temperature=0  # Deterministic for caching
    )
    return response.choices[0].message.content

5. Monitor Usage ​

Track costs in real-time:

python
def log_usage(response):
    usage = response.usage
    print(f"Tokens used: {usage.total_tokens}")
    print(f"Estimated cost: ${calculate_cost(usage)}")

6. Handle Streaming Errors ​

Streaming can fail mid-response:

python
try:
    stream = client.chat.completions.create(..., stream=True)
    for chunk in stream:
        print(chunk.choices[0].delta.content or "", end="")
except Exception as e:
    print(f"\nStream interrupted: {e}")
    # Log partial response, retry, or handle gracefully

Rate Limiting ​

Tokenlio enforces rate limits at multiple levels:

Per-Key Limits ​

Set when creating the key (optional):

  • Requests per minute
  • Requests per day
  • Monthly spend limit

Workspace Limits ​

Default limits per workspace:

  • 60 requests per minute (personal)
  • 120 requests per minute (organization)
  • Contact support for higher limits

Handling 429 Errors ​

python
import time

def handle_rate_limit(error):
    retry_after = int(error.headers.get('Retry-After', 60))
    print(f"Rate limited. Retrying after {retry_after}s")
    time.sleep(retry_after)

Model Compatibility ​

Tokenlio supports OpenAI-compatible parameters for all models:

python
# Works with GPT-4
client.chat.completions.create(
    model="gpt-4-turbo",
    messages=[...],
    temperature=0.7
)

# Also works with Claude
client.chat.completions.create(
    model="claude-3-5-sonnet",
    messages=[...],
    temperature=0.7  # Same API
)

Model-Specific Notes ​

Some parameters may behave differently across models:

  • Check model documentation for specifics
  • Test behavior during development
  • Not all models support all features (e.g., function calling)

Testing ​

Test your integration before production:

python
# Use a small, cheap model for testing
test_response = client.chat.completions.create(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": "test"}],
    max_tokens=5
)

assert test_response.choices[0].message.content
print("✓ API integration working")

Next Steps ​

Need Help? ​

統合インターフェースで主要な AI モデルにアクセス