Streaming
Server-sent events in the OpenAI chunk shape. Set stream: true; the SDKs turn the stream into an async iterator.
Basics
import os from openai import OpenAI client = OpenAI(base_url="https://gridport.ai/v1", api_key=os.environ["GRIDPORT_API_KEY"]) stream = client.chat.completions.create(model="deepseek-ai/DeepSeek-V4-Flash", messages=[{"role": "user", "content": "Count to five."}], stream=True, stream_options={"include_usage": True}) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="") if chunk.usage: # last data chunk before [DONE] print(chunk.usage.total_tokens, chunk.usage.cost)
Each event is a data: line with one JSON chunk; the stream ends with data: [DONE]. The first chunk carries the role, later chunks carry content deltas, and the final chunk has finish_reason. A streamed response carries usage (with cost) in its last data chunk only when the request sets stream_options.include_usage; without it no usage chunk is sent, as with OpenAI. The request log has the usage either way.
Headers arrive first
x-request-id and x-tf-model-version are sent with the 200 status before the first token, so you can log them even if the stream fails later.
Failure after the first byte
Before the first byte the gateway fails over to the next route transparently. After the first byte it cannot: the stream ends with an error event and a final chunk whose finish_reason is "error". You are billed only for the tokens you received; the request shows as partial in the log.
data: {"id":"chatcmpl-…","choices":[{"index":0,"delta":{},"finish_reason":"error"}],"error":{"code":"stream_interrupted","message":"…"}}
data: [DONE]Stopping
Close the connection (AbortController in Node and browsers, closing the response in Python). The gateway cancels the upstream call, the request is logged as cancelled, and only delivered tokens are billed.