Skip to main content
Agentium offers three voice paths. Choose according to who should own the conversation and tool loop. Import voice adapters from @agentium/core/voice. Upgrading from 3.x? Review the v4 voice migration.

Open a native session

Install ws and the optional SDK required by your provider. Set OPENAI_API_KEY for this example. Your application supplies microphone capture and playback.
The connection example does not capture a microphone. Send audio in the negotiated codec/rate and route audio events to your playback device. Use Socket.IO voice gateway for a browser-facing transport or implement a media adapter in your host.

Provider contracts

These are defaults in the checked-out adapter, not guarantees of model access in every account. Pin a model when your deployment requires it. Unsupported provider-specific options reject rather than silently changing semantics. OpenAI temperature and remote reusable prompt objects are rejected. Input transcription defaults to gpt-transcribe for native OpenAI WebSocket sessions. Supply application-owned transcriptionContext for language/keyword hints. See streaming voice for standalone recognizers and synthesizers.

Tools and ownership

Local tools use the normal validation, approval, execution policy, and cancellation path. Provider-executed remote MCP tools cannot be intercepted locally; native remote MCP cannot be combined with local policy/approval configuration. Register MCP functions as ordinary local tools when you need shared enforcement. A memory-enabled VoiceAgent binds to the first connection’s tenant and user. Later owner changes fail, including anonymous-to-authenticated changes. Use separate instances and scoped backing storage per owner. This guard does not partition a shared storage client automatically.

Playback-confirmed history

Generated text is not proof that a user heard it. A session keeps generated speech and playback-confirmed text separate. Route trusted playback offsets through acknowledgePlayback({ generationId, playedCharacters }); confirm completion only after the final transcript has actually played. Interrupted or unacknowledged text remains delivery-unknown in getTranscript(). Native provider-side conversation state is still provider-owned; local playback acknowledgements do not rewrite it. Use AgentVoiceBrain when your application needs coordinator-owned, playback-confirmed history.

Recovery

Recovery is opt-in through VoiceAgentConfig.recovery. Configure total attempts, delays, connection/elapsed deadlines, and fallback: "stop" or an explicit "fresh" fallback.
  • Gemini: session resumption requires a safe idle checkpoint. Input, generation, tool dispatch, or uncertain effects invalidate it.
  • OpenAI: fresh recovery creates a new configured session and loses provider conversation history. It does not replay audio or tools.
  • Either: uncertain effects can prevent recovery. The host must reconcile them before restarting.
Listen for recovery states recovering, recovered, and failed. Pause capture while recovering and request new input after reconnection. See the gateway client contract.

WebRTC and outbound calls

createRealtimeClientSecret and createRealtimeCall are exported from @agentium/core. The latter accepts WebRTC sdp and returns { id, sdp, raw }; the returned SDP is the answer. It rejects sipUri because it does not place outbound SIP calls. Authenticate and authorize your application endpoint before minting a client secret. Outbound carrier call control is a separate telephony API. It does not itself transport audio.

Validation limits

Local contract tests cover tool concurrency, interruption, transcript reconciliation, buffer limits, and cleanup. They do not establish speech accuracy, live carrier interoperability, or latency for your deployment. Measure actual client playback when reporting first-audible latency.