> ## Documentation Index
> Fetch the complete documentation index at: https://docs.agentium.in/llms.txt
> Use this file to discover all available pages before exploring further.

# Choose voice and media adapters

> Find the recognition, synthesis, realtime, media, and phone adapters for your application.

Start with who owns the conversation. A native realtime provider handles speech and reasoning together. A streaming pipeline lets an ordinary Agent use separate recognizers and synthesizers.

<Columns cols={2}>
  <Card title="Native realtime" icon="waveform-lines" href="/voice/overview">VoiceAgent with OpenAI Realtime or Google Gemini Live. Your host connects capture and playback.</Card>
  <Card title="Composable speech pipeline" icon="microphone" href="/voice/streaming">StreamingVoicePipeline with independent STT, Agent reasoning, TTS, and media delivery.</Card>
</Columns>

## Adapter directory

| Adapter | Purpose | Package / optional dependency |
| - | - | - |
| [OpenAIRealtimeProvider](/voice/openai) | Native realtime speech and tools | `@agentium/core/voice`, `ws` |
| [GoogleLiveProvider](/voice/google) | Gemini Live conversation | `@agentium/core/voice`, `@google/genai` |
| [OpenAIStreamingRecognizer](/voice/openai#streaming-recognition) | Dedicated live transcription | `@agentium/core/voice`, `ws` |
| [ElevenLabsRecognizer](/voice/elevenlabs) | Scribe streaming recognition | `@agentium/core/voice`, `ws` |
| [ElevenLabsSynthesizer](/voice/elevenlabs) | Streaming text-to-speech | `@agentium/core/voice`, `ws` |
| [Speech Engine callback](/voice/elevenlabs#speech-engine) | Bring an Agent brain to an authenticated ElevenLabs engine session | `@agentium/core/voice`; host owns engine SDK |
| [SarvamRecognizer](/voice/sarvam) | Indian-language recognition modes | `@agentium/core/voice`, `ws`; preview adapter |
| [SarvamSynthesizer](/voice/sarvam) | Bulbul streaming synthesis | `@agentium/core/voice`, `ws` |
| [LiveKitVoiceTransport](/voice/livekit) | Send PCM frames to a room audio source | `@agentium/core/voice`; host owns `@livekit/rtc-node` |
| [Outbound call providers](/voice/telephony) | Dial, read status, request hangup | `@agentium/core/telephony`; six carrier/SIP adapters |

Speech adapters also accept an injected `socketFactory`; `ws` is required only for their default WebSocket implementation.

## Connect the layers

```mermaid theme={null}
flowchart LR
  Caller[Caller or microphone] --> Media[Host media transport]
  Media --> STT[Speech recognizer]
  STT --> Brain[AgentVoiceBrain]
  Brain --> TTS[Speech synthesizer]
  TTS --> Media
  Calls[OutboundCallService] --> Carrier[Carrier or SIP room]
  Carrier --> Media
```

Your host decodes containers, converts codecs/rates, detects turn boundaries, and confirms playback. Agentium does not infer that a generated or transmitted frame was heard. See [interruption and history](/voice/streaming#history-and-interruption).

## Choose the audio contract

Read each adapter's `capabilities.formats` before opening it. `SpeechFormat` includes encoding, sample rate, and mono channels; every `AudioFrame` also carries turn, generation, and sequence identity. The pipeline does not silently resample audio between incompatible adapters.

For a browser transport, see the [Socket.IO voice gateway](/voice/gateway). For phone applications, begin with [call setup](/telephony/quickstart) and configure the carrier's media endpoint separately. For completed audio files, [VoicePipeline](/voice/overview) offers buffered STT → model → TTS without the Agent tool loop.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.