Skip to content

Streaming Responses ​

Overview ​

Streaming allows you to receive API responses incrementally as they're generated, providing a better user experience for real-time applications like chatbots.

Why Stream? ​

Benefits:

  • ✅ Faster perceived response time
  • ✅ Better UX for long responses
  • ✅ Lower memory usage
  • ✅ Real-time user feedback

Use Cases:

  • Chatbots and conversational AI
  • Real-time content generation
  • Live translation
  • Code generation interfaces

Basic Usage ​

Python ​

python
from openai import OpenAI

client = OpenAI(
    api_key="your-tokenlio-key",
    base_url="https://api.tokenlio.ai/v1"
)

stream = client.chat.completions.create(
    model="gpt-4-turbo",
    messages=[{"role": "user", "content": "Tell me a story"}],
    stream=True  # Enable streaming
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Node.js ​

javascript
import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: 'your-tokenlio-key',
  baseURL: 'https://api.tokenlio.ai/v1'
});

const stream = await client.chat.completions.create({
  model: 'gpt-4-turbo',
  messages: [{ role: 'user', content: 'Tell me a story' }],
  stream: true
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content || '');
}

cURL ​

bash
curl -N https://api.tokenlio.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $TOKENLIO_API_KEY" \
  -d '{
    "model": "gpt-4-turbo",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": true
  }'

Response Format ​

Streaming responses use Server-Sent Events (SSE):

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":null}]}

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-4-turbo","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

Chunk Structure ​

Each chunk contains:

  • id: Unique completion ID
  • object: Always "chat.completion.chunk"
  • created: Unix timestamp
  • model: Model used
  • choices[].delta: Incremental content
  • choices[].finish_reason: Null until last chunk

Finish Reasons ​

  • stop: Natural completion
  • length: Max tokens reached
  • content_filter: Content filtered
  • function_call: Function call completed

Advanced Examples ​

With Function Calling ​

python
tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get current weather",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string"}
                },
                "required": ["location"]
            }
        }
    }
]

stream = client.chat.completions.create(
    model="gpt-4-turbo",
    messages=[{"role": "user", "content": "What's the weather in SF?"}],
    tools=tools,
    stream=True
)

function_name = ""
function_args = ""

for chunk in stream:
    delta = chunk.choices[0].delta
    
    if delta.tool_calls:
        for tool_call in delta.tool_calls:
            if tool_call.function.name:
                function_name = tool_call.function.name
            if tool_call.function.arguments:
                function_args += tool_call.function.arguments

print(f"Function: {function_name}")
print(f"Arguments: {function_args}")

React Component ​

typescript
import { useState } from 'react';
import OpenAI from 'openai';

function ChatStream() {
  const [response, setResponse] = useState('');
  const [loading, setLoading] = useState(false);

  const sendMessage = async (content: string) => {
    setLoading(true);
    setResponse('');

    const client = new OpenAI({
      apiKey: process.env.TOKENLIO_API_KEY,
      baseURL: 'https://api.tokenlio.ai/v1',
      dangerouslyAllowBrowser: true // Only for demo
    });

    const stream = await client.chat.completions.create({
      model: 'gpt-4-turbo',
      messages: [{ role: 'user', content }],
      stream: true
    });

    for await (const chunk of stream) {
      const content = chunk.choices[0]?.delta?.content || '';
      setResponse(prev => prev + content);
    }

    setLoading(false);
  };

  return (
    <div>
      <textarea value={response} readOnly />
      <button onClick={() => sendMessage('Hello!')} disabled={loading}>
        {loading ? 'Generating...' : 'Send'}
      </button>
    </div>
  );
}

Flask API ​

python
from flask import Flask, Response, stream_with_context
from openai import OpenAI

app = Flask(__name__)
client = OpenAI(
    api_key="your-tokenlio-key",
    base_url="https://api.tokenlio.ai/v1"
)

@app.route('/chat', methods=['POST'])
def chat():
    def generate():
        stream = client.chat.completions.create(
            model="gpt-4-turbo",
            messages=[{"role": "user", "content": "Hello"}],
            stream=True
        )
        
        for chunk in stream:
            content = chunk.choices[0].delta.content
            if content:
                yield f"data: {content}\n\n"
        
        yield "data: [DONE]\n\n"
    
    return Response(
        stream_with_context(generate()),
        mimetype='text/event-stream'
    )

Error Handling ​

Connection Errors ​

python
from openai import APIError, APIConnectionError

try:
    stream = client.chat.completions.create(
        model="gpt-4-turbo",
        messages=[{"role": "user", "content": "Hello"}],
        stream=True
    )
    
    for chunk in stream:
        print(chunk.choices[0].delta.content or "", end="")

except APIConnectionError as e:
    print(f"Connection error: {e}")
    # Retry or show error to user

except APIError as e:
    print(f"API error: {e}")

Mid-Stream Errors ​

Errors can occur mid-stream. Handle them gracefully:

python
try:
    for chunk in stream:
        content = chunk.choices[0].delta.content
        if content:
            print(content, end="", flush=True)
except Exception as e:
    print(f"\n[Error: {e}]")
    # Show error indicator to user

Best Practices ​

Buffer Short Chunks ​

For UI updates, buffer very short chunks:

python
buffer = ""
buffer_size = 5  # characters

for chunk in stream:
    content = chunk.choices[0].delta.content or ""
    buffer += content
    
    if len(buffer) >= buffer_size or chunk.choices[0].finish_reason:
        print(buffer, end="", flush=True)
        buffer = ""

Timeout Handling ​

Set longer timeouts for streaming:

python
import httpx

client = OpenAI(
    api_key="tk-...",
    base_url="https://api.tokenlio.ai/v1",
    http_client=httpx.Client(timeout=300.0)  # 5 minutes
)

Rate Limiting ​

Streaming responses count as one request. Token usage is calculated after completion.

Client-Side Streaming ​

Never expose API keys in client-side code. Use a backend proxy:

Client → Your Server → Tokenlio API
         (proxies stream)

Performance Tips ​

  1. Use max_tokens to limit response length and cost
  2. Set temperature lower for faster, more consistent responses
  3. Buffer chunks to reduce UI re-renders
  4. Close streams properly to free resources

Comparison: Stream vs Non-Stream ​

AspectStreamingNon-Streaming
Time to First Token~200msN/A
Total TimeSameSame
Memory UsageLowerHigher
Error HandlingMore complexSimpler
User ExperienceBetter for long responsesFine for short responses
DebuggingHarderEasier

Support ​

Questions about streaming?

통합 인터페이스로 주요 AI 모델에 액세스