@agentium/core/voice: OpenAIRealtimeProvider for a native voice conversation, and OpenAIStreamingRecognizer for transcription in a composable pipeline.
Native realtime
voice.connect({ tenantId, userId, sessionId }), attach audio/transcript callbacks, send negotiated audio, and close the returned session when the conversation ends. The voice session guide shows the connection lifecycle.
V4 defaults to gpt-realtime-2.1, with gpt-transcribe input transcription. It uses nested GA audio configuration, accepts supported PCM/G.711 formats, and rejects unsupported sampling parameters and remote reusable prompts. Do not substitute a text-only model ID.
Streaming recognition
TranscriptEvent values. It accepts gpt-live-transcribe and gpt-transcribe; delay is valid only for the live model. Credentials are resolved when opening a session.
Use open(config, signal), consume session.events concurrently, send monotonically sequenced audio frames, and call flush() at a host-detected turn boundary. Always close the recognizer session. The readiness and committed-turn deadlines bound waiting; they are not audio playback deadlines.
Pass this recognizer to StreamingVoicePipeline with a compatible transport, brain, and synthesizer. For browser credential minting and SDP exchange, use the separately authenticated voice gateway.
OpenAI recovery creates a fresh provider conversation only when explicitly allowed. It does not replay audio or tool results; see recovery.