StreamingVoicePipeline separates speech recognition, reasoning, speech synthesis, and media delivery. AgentVoiceBrain runs an ordinary Agent with its tools and execution controls.
Connect your media transport
This integration fragment requires a host-providedVoiceTransport, an approved ElevenLabs voice ID, and provider keys (SARVAM_API_KEY, ELEVENLABS_API_KEY, and OPENAI_API_KEY). Install the optional ws peer and the chosen model SDK.
AudioFrame values with monotonic sequence numbers to pipeline.sendAudio(frame). flushInput() commits manually segmented input. The host provides resampling, channel mixing, and container decoding when its input format differs from the adapter.
History and interruption
The coordinator supplies canonical history throughRunOpts.history and ephemeral: true. Configure the wrapped Agent without automatic memory; the coordinator owns the conversation. Explicit memory tools can still be registered deliberately.
New input aborts the previous reasoning and synthesis generation. Late output is discarded by generation ID. Only public text is spoken; reasoning and tool payloads are excluded.
Route actual playback offsets to pipeline.acknowledgePlayback(). Without acknowledgements, generated text is marked delivery-unknown. Packet receipt alone does not prove playback.
Available adapters
Check adapter capabilities before selecting codecs/rates. ElevenLabs stream-input synthesis supports its configured v2 endpoint families; v3/v4 are rejected on that endpoint. Speech sockets use bounded queues and buffers; overflow raises
VoiceBackpressureError.
pipeline.timings measures transcript-to-text, first frame offered to transport, and generation duration. These do not measure first-audible playback. Host-owned clients and the Agent require their own shutdown when the host finishes using them.