Chat Completions API
Generate conversational responses using AI models through the Chat Completions endpoint.
Endpoint
POST https://api.tokenlio.ai/v1/chat/completionsAuthentication
Include your API key in the Authorization header:
Authorization: Bearer YOUR_API_KEYRequest Body
Required Parameters
| Parameter | Type | Description |
|---|---|---|
model | string | Model identifier (e.g., gpt-4-turbo) |
messages | array | Array of message objects |
Optional Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
temperature | number | 1.0 | Sampling temperature (0.0-2.0) |
max_tokens | integer | ∞ | Maximum tokens to generate |
top_p | number | 1.0 | Nucleus sampling threshold |
frequency_penalty | number | 0.0 | Penalize token frequency (-2.0 to 2.0) |
presence_penalty | number | 0.0 | Penalize token presence (-2.0 to 2.0) |
stop | string or array | null | Stop sequences |
stream | boolean | false | Stream response tokens |
n | integer | 1 | Number of completions to generate |
user | string | - | Unique identifier for end-user |
Message Format
Each message object contains:
| Field | Type | Required | Description |
|---|---|---|---|
role | string | Yes | One of: system, user, assistant |
content | string | Yes | Message content |
name | string | No | Name of the message author |
Message Roles
system: Sets behavior and context
{"role": "system", "content": "You are a helpful coding assistant."}user: User messages
{"role": "user", "content": "How do I reverse a string in Python?"}assistant: Assistant responses (for conversation history)
{"role": "assistant", "content": "You can use string slicing: text[::-1]"}Examples
Basic Request
curl https://api.tokenlio.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "gpt-4-turbo",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.tokenlio.ai/v1"
)
response = client.chat.completions.create(
model="gpt-4-turbo",
messages=[
{"role": "user", "content": "What is the capital of France?"}
]
)
print(response.choices[0].message.content)import OpenAI from 'openai';
const client = new OpenAI({
apiKey: 'YOUR_API_KEY',
baseURL: 'https://api.tokenlio.ai/v1',
});
const response = await client.chat.completions.create({
model: 'gpt-4-turbo',
messages: [
{ role: 'user', content: 'What is the capital of France?' }
],
});
console.log(response.choices[0].message.content);With System Message
{
"model": "gpt-4-turbo",
"messages": [
{
"role": "system",
"content": "You are a helpful assistant that answers concisely."
},
{
"role": "user",
"content": "Explain quantum computing"
}
],
"max_tokens": 150
}Multi-Turn Conversation
{
"model": "gpt-4-turbo",
"messages": [
{"role": "system", "content": "You are a coding tutor."},
{"role": "user", "content": "How do I sort an array in JavaScript?"},
{"role": "assistant", "content": "You can use the .sort() method."},
{"role": "user", "content": "Can you show me an example?"}
]
}With Temperature Control
# More deterministic (focused)
response = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "List 3 capital cities"}],
temperature=0.2
)
# More creative (varied)
response = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Write a creative story"}],
temperature=1.5
)Response Format
Success Response
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1699999999,
"model": "gpt-4-turbo",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The capital of France is Paris."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 15,
"completion_tokens": 8,
"total_tokens": 23
}
}Response Fields
| Field | Type | Description |
|---|---|---|
id | string | Unique completion ID |
object | string | Object type (chat.completion) |
created | integer | Unix timestamp |
model | string | Model used |
choices | array | Generated completions |
usage | object | Token usage statistics |
Choice Object
| Field | Type | Description |
|---|---|---|
index | integer | Choice index (when n > 1) |
message | object | Generated message |
finish_reason | string | Why generation stopped |
Finish Reasons
stop: Natural completion or stop sequence hitlength: Reachedmax_tokenslimitcontent_filter: Content filtered by safety systemfunction_call: Function call generated (if supported)
Usage Object
| Field | Type | Description |
|---|---|---|
prompt_tokens | integer | Input tokens consumed |
completion_tokens | integer | Output tokens generated |
total_tokens | integer | Total tokens (prompt + completion) |
Streaming Responses
Stream tokens as they're generated for real-time output.
Enable Streaming
Set stream: true in the request:
{
"model": "gpt-4-turbo",
"messages": [...],
"stream": true
}Streaming Format
Responses are sent as Server-Sent Events (SSE):
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1699999999,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1699999999,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1699999999,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1699999999,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]Python Streaming Example
stream = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Tell me a story"}],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)JavaScript Streaming Example
const stream = await client.chat.completions.create({
model: 'gpt-4-turbo',
messages: [{ role: 'user', content: 'Tell me a story' }],
stream: true,
});
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content || '';
process.stdout.write(content);
}Advanced Features
Stop Sequences
Stop generation when specific strings are encountered:
{
"model": "gpt-4-turbo",
"messages": [...],
"stop": ["\n\n", "END", "---"]
}Multiple Completions
Generate multiple responses (uses more tokens):
{
"model": "gpt-4-turbo",
"messages": [...],
"n": 3
}Response includes 3 choices:
{
"choices": [
{"index": 0, "message": {"content": "Response 1"}},
{"index": 1, "message": {"content": "Response 2"}},
{"index": 2, "message": {"content": "Response 3"}}
]
}Penalties
Control repetition and diversity:
response = client.chat.completions.create(
model="gpt-4-turbo",
messages=[...],
frequency_penalty=0.5, # Reduce repetition
presence_penalty=0.5 # Encourage new topics
)Error Handling
Error Response Format
{
"error": {
"message": "Invalid API key provided",
"type": "invalid_request_error",
"code": "invalid_api_key"
}
}Common Errors
| Status | Error Type | Description | Solution |
|---|---|---|---|
| 400 | invalid_request_error | Malformed request | Check request format |
| 401 | authentication_error | Invalid API key | Verify API key |
| 402 | insufficient_balance | No balance | Top up account |
| 403 | permission_error | Model not accessible | Check model access |
| 429 | rate_limit_error | Too many requests | Implement backoff |
| 500 | api_error | Server error | Retry request |
| 503 | service_unavailable | Service down | Try again later |
Error Handling Example
from openai import OpenAI, OpenAIError
try:
response = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Hello"}]
)
except OpenAIError as e:
print(f"Error: {e}")
# Log error, retry, or handle gracefullyBest Practices
1. Use System Messages
Set clear behavior guidelines:
{
"messages": [
{"role": "system", "content": "You are a helpful, accurate assistant. Be concise."},
{"role": "user", "content": "..."}
]
}2. Limit Output Length
Control costs with max_tokens:
{
"max_tokens": 500
}3. Optimize Context
Only include necessary conversation history:
# Keep last 10 messages
recent_messages = conversation[-10:]4. Handle Errors Gracefully
Implement retry logic and user feedback:
try:
response = client.chat.completions.create(...)
except RateLimitError:
# Wait and retry
time.sleep(5)
response = client.chat.completions.create(...)
except Exception as e:
# Show user-friendly error
return "Sorry, something went wrong. Please try again."5. Monitor Token Usage
Track usage in responses:
usage = response.usage
print(f"Tokens used: {usage.total_tokens}")
cost = calculate_cost(usage)
print(f"Cost: ${cost:.4f}")Rate Limits
Rate limits depend on your workspace and key settings:
- Default: 60 requests/minute
- Organization: 120 requests/minute
- Custom: Contact support for higher limits
Implement exponential backoff when hitting limits:
import time
def make_request_with_backoff(max_retries=3):
for i in range(max_retries):
try:
return client.chat.completions.create(...)
except RateLimitError:
if i == max_retries - 1:
raise
time.sleep(2 ** i) # 1s, 2s, 4sModel Compatibility
All models support the chat completions format:
# GPT-4
client.chat.completions.create(model="gpt-4-turbo", ...)
# Claude
client.chat.completions.create(model="claude-3-5-sonnet", ...)
# Gemini
client.chat.completions.create(model="gemini-1.5-pro", ...)See available models for the complete list.