流式响应 (Streaming)
在聊天界面或实时协作场景中,使用流式响应(Server-Sent Events / SSE)可以显著降低首字延迟 (TTFT),提供打字机式的流畅用户体验。
启用流式传输
只需在调用参数中指定 "stream": true:
Python 示例
python
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("TOKENLIO_API_KEY"),
base_url="https://api.tokenlio.ai/v1"
)
stream = client.chat.completions.create(
model="claude-3-5-sonnet",
messages=[{"role": "user", "content": "写一首赞美科技与智能的十四行诗"}],
stream=True
)
for chunk in stream:
delta = chunk.choices[0].delta.content or ""
print(delta, end="", flush=True)
print()Node.js 示例
javascript
import OpenAI from 'openai';
const client = new OpenAI({
baseURL: 'https://api.tokenlio.ai/v1',
apiKey: process.env.TOKENLIO_API_KEY,
});
const stream = await client.chat.completions.create({
model: 'deepseek-chat',
messages: [{ role: 'user', content: '分析微服务高可用方案' }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content || '');
}数据块格式 (SSE)
data: {"id":"chatcmpl-xxx","object":"chat.completion.chunk","choices":[{"delta":{"content":"你好"},"index":0,"finish_reason":null}]}
data: {"id":"chatcmpl-xxx","object":"chat.completion.chunk","choices":[{"delta":{"content":"!"},"index":0,"finish_reason":null}]}
data: [DONE]