Docs menu

Docs

Streaming

Server-sent events, chunk by chunk, with usage in the final frame.

Set "stream": true and the response becomes a stream of server-sent events. Chunks arrive as the model generates, which is how you build responsive UIs and keep long answers feeling instant.

#The wire format

Each event is a data: line carrying a JSON chunk. Deltas accumulate in order: concatenate choices[0].delta.content to rebuild the answer. A final chunk reports usage, and the stream ends with a data: [DONE] sentinel.

curl
curl -N https://api.dipoleml.com/v1/chat/completions \
  -H "Authorization: Bearer $RAVEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Raven Max",
    "stream": true,
    "messages": [{"role": "user", "content": "Reply with exactly: STREAM-OK"}]
  }'
event stream (abridged)
data: {"id":"gen-1791206856-JD71sb6wPUzNlMfNGOTI","object":"chat.completion.chunk","model":"Raven Max","choices":[{"index":0,"delta":{"content":"STREAM-OK","role":"assistant"},"finish_reason":null}]}

data: {"id":"gen-1791206856-JD71sb6wPUzNlMfNGOTI","object":"chat.completion.chunk","model":"Raven Max","choices":[{"index":0,"delta":{"content":""},"finish_reason":"stop"}]}

data: {"id":"gen-1791206856-JD71sb6wPUzNlMfNGOTI","object":"chat.completion.chunk","model":"Raven Max","choices":[],"usage":{"prompt_tokens":608,"completion_tokens":5,"total_tokens":613}}

data: [DONE]

Three things to know about the stream:

  • Usage arrives in the final content chunk. You do not need to request it; the service always appends a usage chunk with empty choices.
  • Reasoning may stream too. Some chunks carry deltas in a reasoning or reasoning_content field instead of content. Treat it as plain text; see Reasoning.
  • Errors mid-stream close the connection. If a failure happens after tokens were sent, the stream ends early. Retry only if no content was received.

#Python

python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.dipoleml.com/v1",
    api_key=os.environ["RAVEN_API_KEY"],
)

stream = client.chat.completions.create(
    model="Raven Flash",
    stream=True,
    messages=[{"role": "user", "content": "Explain caching in three sentences"}],
)
for chunk in stream:
    delta = chunk.choices[0].delta
    if delta.content:
        print(delta.content, end="", flush=True)

#Retry guidance

Retry a failed stream only when nothing was delivered yet. A safe pattern: attempt the request, and if the connection fails or returns an error before the first content delta, retry with exponential backoff up to three times. If content already arrived, surface an honest message and let the user resend. The raven app follows exactly this pattern.