Skip to main content
StreamingVoicePipeline separates speech recognition, reasoning, speech synthesis, and media delivery. AgentVoiceBrain runs an ordinary Agent with its tools and execution controls.

Connect your media transport

This integration fragment requires a host-provided VoiceTransport, an approved ElevenLabs voice ID, and provider keys (SARVAM_API_KEY, ELEVENLABS_API_KEY, and OPENAI_API_KEY). Install the optional ws peer and the chosen model SDK.
Supply negotiated AudioFrame values with monotonic sequence numbers to pipeline.sendAudio(frame). flushInput() commits manually segmented input. The host provides resampling, channel mixing, and container decoding when its input format differs from the adapter.

History and interruption

The coordinator supplies canonical history through RunOpts.history and ephemeral: true. Configure the wrapped Agent without automatic memory; the coordinator owns the conversation. Explicit memory tools can still be registered deliberately. New input aborts the previous reasoning and synthesis generation. Late output is discarded by generation ID. Only public text is spoken; reasoning and tool payloads are excluded. Route actual playback offsets to pipeline.acknowledgePlayback(). Without acknowledgements, generated text is marked delivery-unknown. Packet receipt alone does not prove playback.

Available adapters

Check adapter capabilities before selecting codecs/rates. ElevenLabs stream-input synthesis supports its configured v2 endpoint families; v3/v4 are rejected on that endpoint. Speech sockets use bounded queues and buffers; overflow raises VoiceBackpressureError. pipeline.timings measures transcript-to-text, first frame offered to transport, and generation duration. These do not measure first-audible playback. Host-owned clients and the Agent require their own shutdown when the host finishes using them.