Native realtime
VoiceAgent with OpenAI Realtime or Google Gemini Live. Your host connects capture and playback.
Composable speech pipeline
StreamingVoicePipeline with independent STT, Agent reasoning, TTS, and media delivery.
Adapter directory
Speech adapters also accept an injected
socketFactory; ws is required only for their default WebSocket implementation.
Connect the layers
Your host decodes containers, converts codecs/rates, detects turn boundaries, and confirms playback. Agentium does not infer that a generated or transmitted frame was heard. See interruption and history.Choose the audio contract
Read each adapter’scapabilities.formats before opening it. SpeechFormat includes encoding, sample rate, and mono channels; every AudioFrame also carries turn, generation, and sequence identity. The pipeline does not silently resample audio between incompatible adapters.
For a browser transport, see the Socket.IO voice gateway. For phone applications, begin with call setup and configure the carrier’s media endpoint separately. For completed audio files, VoicePipeline offers buffered STT → model → TTS without the Agent tool loop.