Skip to main content
Agentium exposes two different OpenAI voice adapters from @agentium/core/voice: OpenAIRealtimeProvider for a native voice conversation, and OpenAIStreamingRecognizer for transcription in a composable pipeline.

Native realtime

Construction does not connect. Call voice.connect({ tenantId, userId, sessionId }), attach audio/transcript callbacks, send negotiated audio, and close the returned session when the conversation ends. The voice session guide shows the connection lifecycle. V4 defaults to gpt-realtime-2.1, with gpt-transcribe input transcription. It uses nested GA audio configuration, accepts supported PCM/G.711 formats, and rejects unsupported sampling parameters and remote reusable prompts. Do not substitute a text-only model ID.

Streaming recognition

This recognizer accepts mono PCM16 at 24 kHz and emits partial/final TranscriptEvent values. It accepts gpt-live-transcribe and gpt-transcribe; delay is valid only for the live model. Credentials are resolved when opening a session. Use open(config, signal), consume session.events concurrently, send monotonically sequenced audio frames, and call flush() at a host-detected turn boundary. Always close the recognizer session. The readiness and committed-turn deadlines bound waiting; they are not audio playback deadlines. Pass this recognizer to StreamingVoicePipeline with a compatible transport, brain, and synthesizer. For browser credential minting and SDP exchange, use the separately authenticated voice gateway. OpenAI recovery creates a fresh provider conversation only when explicitly allowed. It does not replay audio or tool results; see recovery.