Mozaik

Streaming

Pass streaming true on InferenceInput to publish each provider chunk as inference.stream.

Pass streaming: true on the InferenceInput you give runLoop. The loop takes the inference_streaming path and publishes each provider chunk as inference.stream. The payload is the inner provider event (type + payload).

runLoop(agent.getId(), message, {
  model: 'gpt-5.5',
  context: agent.getMemory().getContext(),
  tools: agent.getTools(),
  streaming: true,
});

Requesting streaming for a model whose specification has supportsStreaming: false fails validation before the API is called.

Handle chunks with a situation handler that matches inference.stream:

import { SituationSpecification, type SituationHandler, type SituationContext } from '@mozaik-ai/core';

class WhenInferenceStreams extends SituationSpecification {
  isSatisfiedBy({ event }: SituationContext): boolean {
    return event.type === 'inference.stream';
  }
}

const liveUi: SituationHandler = {
  specification: new WhenInferenceStreams(),
  processor: {
    apply({ event }) {
      const inner = event.payload as { type?: string; payload?: { delta?: string } };
      if (inner.type === 'response.output_text.delta') {
        process.stdout.write(String(inner.payload?.delta ?? ''));
      }
    },
  },
};

When streaming is off, the loop stays on inference and only publishes inference.started / inference.completed. Completed model output still arrives as inference.completed (and then model.answer) in both modes.

inference.completed carries InferenceOutput: items (FunctionCallItem | ReasoningItem | ModelMessageItem), tokenUsage, and rowResponse.