Nirmos

Streaming

Consume OpenAI-compatible text and tool-call deltas as they arrive.

Both clients consume the gateway's Server-Sent Events stream.

const stream = await nirmos.gateway.chat.stream({
  model: "openai/gpt-5.4-mini",
  messages: [{ role: "user", content: "Write a short release note." }],
});

for await (const event of stream) {
  if (event.type === "content.delta") {
    process.stdout.write(event.delta ?? "");
  }
}

const completion = await stream.finalResponse();
console.log(completion.usage);
const stream = await client.chat.completions.create({
  model: "openai/gpt-5.4-mini",
  messages: [{ role: "user", content: "Write a short release note." }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta.content ?? "");
}

The Nirmos SDK parses SSE boundaries and assembles a normalized final completion. The OpenAI SDK exposes standard chat completion chunks.

Nirmos stream events

EventMeaning
content.deltaNew text for one choice
toolCall.deltaPartial function tool call
completion.completedFully assembled normalized completion

Event names are extensible. Ignore events the application does not use.

for await (const event of stream) {
  switch (event.type) {
    case "content.delta":
      process.stdout.write(event.delta ?? "");
      break;
    case "toolCall.delta":
      consumeToolCallDelta(event.toolCall);
      break;
    case "completion.completed":
      recordUsage(event.completion?.usage);
      break;
  }
}

Final response and one-consumer rule

After iteration finishes, finalResponse() returns the same ChatCompletion shape as chat.create. If you only need complete text and do not need live deltas, call await stream.text() without separately iterating.

A ChatStream can be consumed once. Do not iterate from multiple tasks, and do not call finalResponse() while another task is still iterating it.

Routes

const stream = await nirmos.gateway.chat.stream({
  route: "production-chat",
  messages,
});
const stream = await client.chat.completions.create({
  model: "nirmos/production-chat",
  messages,
  stream: true,
});

Cancellation

Pass request options as the second Nirmos SDK argument:

const stream = await nirmos.gateway.chat.stream(
  { model: "openai/gpt-5.4-mini", messages },
  { signal: request.signal, timeoutMs: 15_000 },
);

Caller cancellation raises RequestAbortedError. Timeouts raise TimeoutError; malformed or server-declared stream failures raise StreamError.

On this page