Skip to content

流式响应 (Streaming) ​

在聊天界面或实时协作场景中,使用流式响应(Server-Sent Events / SSE)可以显著降低首字延迟 (TTFT),提供打字机式的流畅用户体验。

启用流式传输 ​

只需在调用参数中指定 "stream": true:

Python 示例 ​

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ.get("TOKENLIO_API_KEY"),
    base_url="https://api.tokenlio.ai/v1"
)

stream = client.chat.completions.create(
    model="claude-3-5-sonnet",
    messages=[{"role": "user", "content": "写一首赞美科技与智能的十四行诗"}],
    stream=True
)

for chunk in stream:
    delta = chunk.choices[0].delta.content or ""
    print(delta, end="", flush=True)
print()

Node.js 示例 ​

javascript
import OpenAI from 'openai';

const client = new OpenAI({
  baseURL: 'https://api.tokenlio.ai/v1',
  apiKey: process.env.TOKENLIO_API_KEY,
});

const stream = await client.chat.completions.create({
  model: 'deepseek-chat',
  messages: [{ role: 'user', content: '分析微服务高可用方案' }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content || '');
}

数据块格式 (SSE) ​

data: {"id":"chatcmpl-xxx","object":"chat.completion.chunk","choices":[{"delta":{"content":"你好"},"index":0,"finish_reason":null}]}

data: {"id":"chatcmpl-xxx","object":"chat.completion.chunk","choices":[{"delta":{"content":"!"},"index":0,"finish_reason":null}]}

data: [DONE]

通过统一接口访问领先的 AI 模型