Streaming
Pass streaming true on InferenceInput to publish each provider chunk as inference.stream.
Pass streaming: true on the InferenceInput you give runLoop. The loop takes the inference_streaming path and publishes each provider chunk as inference.stream. The payload is the inner provider event (type + payload).
runLoop(agent.getId(), message, {
model: 'gpt-5.5',
context: agent.getMemory().getContext(),
tools: agent.getTools(),
streaming: true,
});Requesting streaming for a model whose specification has supportsStreaming: false fails validation before the API is called.
Handle chunks with a situation handler that matches inference.stream:
import { SituationSpecification, type SituationHandler, type SituationContext } from '@mozaik-ai/core';
class WhenInferenceStreams extends SituationSpecification {
isSatisfiedBy({ event }: SituationContext): boolean {
return event.type === 'inference.stream';
}
}
const liveUi: SituationHandler = {
specification: new WhenInferenceStreams(),
processor: {
apply({ event }) {
const inner = event.payload as { type?: string; payload?: { delta?: string } };
if (inner.type === 'response.output_text.delta') {
process.stdout.write(String(inner.payload?.delta ?? ''));
}
},
},
};When streaming is off, the loop stays on inference and only publishes inference.started / inference.completed. Completed model output still arrives as inference.completed (and then model.answer) in both modes.
inference.completed carries InferenceOutput: items (FunctionCallItem | ReasoningItem | ModelMessageItem), tokenUsage, and rowResponse.