Skip to main content
The recognition and synthesis adapters use WebSockets and can be selected independently. Both import from @agentium/core/voice.

Recognize and synthesize

Supply a voice ID your account may use. The stream-input adapter rejects v3/v4 model families because they require a different dialogue API. A newer model’s existence does not make it compatible with this endpoint. For direct use, open a session with SpeechOpenConfig and an abort signal, consume its events/frames concurrently, send input, then flush and close. A synthesizer’s flush() finishes that generation. To interrupt, abort/close it and start a new generation; do not append to a flushed session. See streaming voice for composing these with AgentVoiceBrain and a media transport.

Speech Engine

Use createElevenLabsSpeechEngineHandler when ElevenLabs owns the audio session and you supply the text reasoning component.
Pass the resulting callback to the official SDK’s authenticated engine.attach(...) integration in your host. Agentium does not create the engine server or verify that connection. Keep SDK authentication enabled. The callback accepts the engine transcript as authoritative history, avoids identical duplicate transcript requests for a session, and streams public text deltas. Do not append the same transcript again into another automatic memory store. The engine/Agent and any host SDK clients retain their own lifecycle.