Skip to main content
Start with who owns the conversation. A native realtime provider handles speech and reasoning together. A streaming pipeline lets an ordinary Agent use separate recognizers and synthesizers.

Native realtime

VoiceAgent with OpenAI Realtime or Google Gemini Live. Your host connects capture and playback.

Composable speech pipeline

StreamingVoicePipeline with independent STT, Agent reasoning, TTS, and media delivery.

Adapter directory

Speech adapters also accept an injected socketFactory; ws is required only for their default WebSocket implementation.

Connect the layers

Your host decodes containers, converts codecs/rates, detects turn boundaries, and confirms playback. Agentium does not infer that a generated or transmitted frame was heard. See interruption and history.

Choose the audio contract

Read each adapter’s capabilities.formats before opening it. SpeechFormat includes encoding, sample rate, and mono channels; every AudioFrame also carries turn, generation, and sequence identity. The pipeline does not silently resample audio between incompatible adapters. For a browser transport, see the Socket.IO voice gateway. For phone applications, begin with call setup and configure the carrier’s media endpoint separately. For completed audio files, VoicePipeline offers buffered STT → model → TTS without the Agent tool loop.