Making Requests
Learn how to make API requests to Tokenlio, handle responses, and implement best practices for production applications.
Base URL
All API requests should use:
https://api.tokenlio.ai/v1Request Format
Tokenlio follows the OpenAI API specification. If you're familiar with OpenAI's API, you're already familiar with Tokenlio.
Basic Request Structure
curl https://api.tokenlio.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "gpt-4-turbo",
"messages": [
{"role": "user", "content": "Hello!"}
]
}'Required Headers
Content-Type: application/json- JSON request bodyAuthorization: Bearer YOUR_API_KEY- Your API key
Response Format
Success Response
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1699999999,
"model": "gpt-4-turbo",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I assist you today?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 10,
"completion_tokens": 12,
"total_tokens": 22
}
}Error Response
{
"error": {
"message": "Insufficient balance",
"type": "insufficient_balance",
"code": "insufficient_balance"
}
}Common Parameters
Model Selection
Specify which model to use:
{
"model": "gpt-4-turbo"
}See available models for the full list.
Messages Array
Chat completions use a messages array:
{
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is AI?"},
{"role": "assistant", "content": "AI stands for..."},
{"role": "user", "content": "Tell me more"}
]
}Roles:
system: Sets assistant behavior (optional, but recommended)user: User messagesassistant: Previous assistant responses (for context)
Temperature
Controls randomness (0.0 to 2.0):
{
"temperature": 0.7
}0.0: Deterministic, focused0.7: Balanced (default)1.0+: More creative, random
Max Tokens
Limit output length:
{
"max_tokens": 500
}TIP
Set max_tokens to control costs. The model stops generating when this limit is reached.
Other Parameters
{
"top_p": 0.9, // Alternative to temperature
"frequency_penalty": 0, // Reduce repetition (-2.0 to 2.0)
"presence_penalty": 0, // Encourage new topics (-2.0 to 2.0)
"stop": ["\n", "END"] // Stop sequences
}Streaming Responses
Get tokens as they're generated instead of waiting for completion.
Python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.tokenlio.ai/v1"
)
stream = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Tell me a story"}],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)JavaScript
const stream = await client.chat.completions.create({
model: 'gpt-4-turbo',
messages: [{ role: 'user', content: 'Tell me a story' }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content || '');
}cURL (Server-Sent Events)
curl https://api.tokenlio.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "gpt-4-turbo",
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}'Error Handling
Status Codes
| Code | Meaning | Common Causes |
|---|---|---|
| 200 | Success | Request completed |
| 400 | Bad Request | Invalid parameters |
| 401 | Unauthorized | Invalid API key |
| 402 | Payment Required | Insufficient balance |
| 403 | Forbidden | Model not accessible |
| 429 | Rate Limited | Too many requests |
| 500 | Server Error | Internal error |
| 503 | Service Unavailable | Temporary outage |
Python Error Handling
from openai import OpenAI, OpenAIError, RateLimitError, AuthenticationError
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.tokenlio.ai/v1"
)
try:
response = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Hello"}]
)
except AuthenticationError:
print("Invalid API key")
except RateLimitError:
print("Rate limit exceeded - slow down")
except OpenAIError as e:
print(f"API error: {e}")JavaScript Error Handling
try {
const response = await client.chat.completions.create({
model: 'gpt-4-turbo',
messages: [{ role: 'user', content: 'Hello' }],
});
} catch (error) {
if (error.status === 401) {
console.error('Invalid API key');
} else if (error.status === 429) {
console.error('Rate limit exceeded');
} else {
console.error('API error:', error);
}
}Retry Logic
Implement exponential backoff for transient errors:
import time
from openai import OpenAI, RateLimitError
def make_request_with_retry(client, max_retries=3):
for attempt in range(max_retries):
try:
return client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Hello"}]
)
except RateLimitError:
if attempt == max_retries - 1:
raise
wait_time = (2 ** attempt) # 1s, 2s, 4s
print(f"Rate limited. Waiting {wait_time}s...")
time.sleep(wait_time)Request IDs
Every response includes a unique ID:
{
"id": "chatcmpl-abc123",
...
}Use this ID when contacting support - it helps us debug issues quickly.
Best Practices
1. Set Timeouts
Prevent hanging requests:
import httpx
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.tokenlio.ai/v1",
timeout=30.0, # 30 second timeout
http_client=httpx.Client()
)2. Use System Messages
Guide model behavior:
{
"messages": [
{
"role": "system",
"content": "You are a technical support assistant. Be concise and helpful."
},
{"role": "user", "content": "How do I reset my password?"}
]
}3. Validate Input
Check user input before sending:
def validate_message(content: str) -> bool:
if not content or not content.strip():
return False
if len(content) > 10000: # Max length
return False
return True4. Cache Responses
Avoid duplicate API calls:
from functools import lru_cache
@lru_cache(maxsize=100)
def get_completion(prompt: str) -> str:
response = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": prompt}],
temperature=0 # Deterministic for caching
)
return response.choices[0].message.content5. Monitor Usage
Track costs in real-time:
def log_usage(response):
usage = response.usage
print(f"Tokens used: {usage.total_tokens}")
print(f"Estimated cost: ${calculate_cost(usage)}")6. Handle Streaming Errors
Streaming can fail mid-response:
try:
stream = client.chat.completions.create(..., stream=True)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
except Exception as e:
print(f"\nStream interrupted: {e}")
# Log partial response, retry, or handle gracefullyRate Limiting
Tokenlio enforces rate limits at multiple levels:
Per-Key Limits
Set when creating the key (optional):
- Requests per minute
- Requests per day
- Monthly spend limit
Workspace Limits
Default limits per workspace:
- 60 requests per minute (personal)
- 120 requests per minute (organization)
- Contact support for higher limits
Handling 429 Errors
import time
def handle_rate_limit(error):
retry_after = int(error.headers.get('Retry-After', 60))
print(f"Rate limited. Retrying after {retry_after}s")
time.sleep(retry_after)Model Compatibility
Tokenlio supports OpenAI-compatible parameters for all models:
# Works with GPT-4
client.chat.completions.create(
model="gpt-4-turbo",
messages=[...],
temperature=0.7
)
# Also works with Claude
client.chat.completions.create(
model="claude-3-5-sonnet",
messages=[...],
temperature=0.7 # Same API
)Model-Specific Notes
Some parameters may behave differently across models:
- Check model documentation for specifics
- Test behavior during development
- Not all models support all features (e.g., function calling)
Testing
Test your integration before production:
# Use a small, cheap model for testing
test_response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "test"}],
max_tokens=5
)
assert test_response.choices[0].message.content
print("✓ API integration working")Next Steps
- Explore available models and pricing
- Learn about billing
- Read the Chat API reference
- Review authentication details
