Streaming Responses
Overview
Streaming allows you to receive API responses incrementally as they're generated, providing a better user experience for real-time applications like chatbots.
Why Stream?
Benefits:
- ✅ Faster perceived response time
- ✅ Better UX for long responses
- ✅ Lower memory usage
- ✅ Real-time user feedback
Use Cases:
- Chatbots and conversational AI
- Real-time content generation
- Live translation
- Code generation interfaces
Basic Usage
Python
python
from openai import OpenAI
client = OpenAI(
api_key="your-tokenlio-key",
base_url="https://api.tokenlio.ai/v1"
)
stream = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Tell me a story"}],
stream=True # Enable streaming
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Node.js
javascript
import OpenAI from 'openai';
const client = new OpenAI({
apiKey: 'your-tokenlio-key',
baseURL: 'https://api.tokenlio.ai/v1'
});
const stream = await client.chat.completions.create({
model: 'gpt-4-turbo',
messages: [{ role: 'user', content: 'Tell me a story' }],
stream: true
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content || '');
}cURL
bash
curl -N https://api.tokenlio.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $TOKENLIO_API_KEY" \
-d '{
"model": "gpt-4-turbo",
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}'Response Format
Streaming responses use Server-Sent Events (SSE):
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]Chunk Structure
Each chunk contains:
id: Unique completion IDobject: Always"chat.completion.chunk"created: Unix timestampmodel: Model usedchoices[].delta: Incremental contentchoices[].finish_reason: Null until last chunk
Finish Reasons
stop: Natural completionlength: Max tokens reachedcontent_filter: Content filteredfunction_call: Function call completed
Advanced Examples
With Function Calling
python
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"}
},
"required": ["location"]
}
}
}
]
stream = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "What's the weather in SF?"}],
tools=tools,
stream=True
)
function_name = ""
function_args = ""
for chunk in stream:
delta = chunk.choices[0].delta
if delta.tool_calls:
for tool_call in delta.tool_calls:
if tool_call.function.name:
function_name = tool_call.function.name
if tool_call.function.arguments:
function_args += tool_call.function.arguments
print(f"Function: {function_name}")
print(f"Arguments: {function_args}")React Component
typescript
import { useState } from 'react';
import OpenAI from 'openai';
function ChatStream() {
const [response, setResponse] = useState('');
const [loading, setLoading] = useState(false);
const sendMessage = async (content: string) => {
setLoading(true);
setResponse('');
const client = new OpenAI({
apiKey: process.env.TOKENLIO_API_KEY,
baseURL: 'https://api.tokenlio.ai/v1',
dangerouslyAllowBrowser: true // Only for demo
});
const stream = await client.chat.completions.create({
model: 'gpt-4-turbo',
messages: [{ role: 'user', content }],
stream: true
});
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content || '';
setResponse(prev => prev + content);
}
setLoading(false);
};
return (
<div>
<textarea value={response} readOnly />
<button onClick={() => sendMessage('Hello!')} disabled={loading}>
{loading ? 'Generating...' : 'Send'}
</button>
</div>
);
}Flask API
python
from flask import Flask, Response, stream_with_context
from openai import OpenAI
app = Flask(__name__)
client = OpenAI(
api_key="your-tokenlio-key",
base_url="https://api.tokenlio.ai/v1"
)
@app.route('/chat', methods=['POST'])
def chat():
def generate():
stream = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Hello"}],
stream=True
)
for chunk in stream:
content = chunk.choices[0].delta.content
if content:
yield f"data: {content}\n\n"
yield "data: [DONE]\n\n"
return Response(
stream_with_context(generate()),
mimetype='text/event-stream'
)Error Handling
Connection Errors
python
from openai import APIError, APIConnectionError
try:
stream = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Hello"}],
stream=True
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
except APIConnectionError as e:
print(f"Connection error: {e}")
# Retry or show error to user
except APIError as e:
print(f"API error: {e}")Mid-Stream Errors
Errors can occur mid-stream. Handle them gracefully:
python
try:
for chunk in stream:
content = chunk.choices[0].delta.content
if content:
print(content, end="", flush=True)
except Exception as e:
print(f"\n[Error: {e}]")
# Show error indicator to userBest Practices
Buffer Short Chunks
For UI updates, buffer very short chunks:
python
buffer = ""
buffer_size = 5 # characters
for chunk in stream:
content = chunk.choices[0].delta.content or ""
buffer += content
if len(buffer) >= buffer_size or chunk.choices[0].finish_reason:
print(buffer, end="", flush=True)
buffer = ""Timeout Handling
Set longer timeouts for streaming:
python
import httpx
client = OpenAI(
api_key="tk-...",
base_url="https://api.tokenlio.ai/v1",
http_client=httpx.Client(timeout=300.0) # 5 minutes
)Rate Limiting
Streaming responses count as one request. Token usage is calculated after completion.
Client-Side Streaming
Never expose API keys in client-side code. Use a backend proxy:
Client → Your Server → Tokenlio API
(proxies stream)Performance Tips
- Use
max_tokensto limit response length and cost - Set
temperaturelower for faster, more consistent responses - Buffer chunks to reduce UI re-renders
- Close streams properly to free resources
Comparison: Stream vs Non-Stream
| Aspect | Streaming | Non-Streaming |
|---|---|---|
| Time to First Token | ~200ms | N/A |
| Total Time | Same | Same |
| Memory Usage | Lower | Higher |
| Error Handling | More complex | Simpler |
| User Experience | Better for long responses | Fine for short responses |
| Debugging | Harder | Easier |
Support
Questions about streaming?
- Email: support@tokenlio.ai
- Documentation: /docs/
