> ## Documentation Index
> Fetch the complete documentation index at: https://docs.agentium.in/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming voice with Agent

> Connect recognizers, an Agent tool loop, synthesizers, and a host media transport.

`StreamingVoicePipeline` separates speech recognition, reasoning, speech synthesis, and media delivery. `AgentVoiceBrain` runs an ordinary Agent with its tools and execution controls.

## Connect your media transport

This integration fragment requires a host-provided `VoiceTransport`, an approved ElevenLabs voice ID, and provider keys (`SARVAM_API_KEY`, `ELEVENLABS_API_KEY`, and `OPENAI_API_KEY`). Install the optional `ws` peer and the chosen model SDK.

```typescript theme={null}
import { Agent, openai } from "@agentium/core";
import {
  AgentVoiceBrain,
  ElevenLabsSynthesizer,
  SarvamRecognizer,
  StreamingVoicePipeline,
  type VoiceTransport,
} from "@agentium/core/voice";

async function openVoice(transport: VoiceTransport, voiceId: string) {
  const agent = new Agent({ name: "support", model: openai(process.env.OPENAI_MODEL ?? "gpt-6.1-sol") });
  const pipeline = new StreamingVoicePipeline({
    recognizer: new SarvamRecognizer(),
    synthesizer: new ElevenLabsSynthesizer({ voiceId }),
    brain: new AgentVoiceBrain(agent, { tenantId: "local-demo", userId: "developer" }),
    transport,
    language: "hi-IN",
  });
  try {
    await pipeline.open();
  } catch (error) {
    await pipeline.close();
    await agent.close();
    throw error;
  }
  return { pipeline, agent }; // The host closes both when the call ends.
}
```

Supply negotiated `AudioFrame` values with monotonic sequence numbers to `pipeline.sendAudio(frame)`. `flushInput()` commits manually segmented input. The host provides resampling, channel mixing, and container decoding when its input format differs from the adapter.

## History and interruption

The coordinator supplies canonical history through `RunOpts.history` and `ephemeral: true`. Configure the wrapped Agent without automatic memory; the coordinator owns the conversation. Explicit memory tools can still be registered deliberately.

New input aborts the previous reasoning and synthesis generation. Late output is discarded by generation ID. Only public text is spoken; reasoning and tool payloads are excluded.

Route actual playback offsets to `pipeline.acknowledgePlayback()`. Without acknowledgements, generated text is marked delivery-unknown. Packet receipt alone does not prove playback.

## Available adapters

| Capability | Adapters |
| - | - |
| Recognition | `OpenAIStreamingRecognizer`, `ElevenLabsRecognizer`, `SarvamRecognizer` |
| Synthesis | `ElevenLabsSynthesizer`, `SarvamSynthesizer` |
| Agent reasoning | `AgentVoiceBrain` |
| Media | Host `VoiceTransport` or optional `LiveKitVoiceTransport` |
| Speech Engine callback | `createElevenLabsSpeechEngineHandler` |

Check adapter capabilities before selecting codecs/rates. ElevenLabs stream-input synthesis supports its configured v2 endpoint families; v3/v4 are rejected on that endpoint. Speech sockets use bounded queues and buffers; overflow raises `VoiceBackpressureError`.

`pipeline.timings` measures transcript-to-text, first frame offered to transport, and generation duration. These do not measure first-audible playback. Host-owned clients and the Agent require their own shutdown when the host finishes using them.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.