Skip to main content
Import SarvamRecognizer and SarvamSynthesizer from @agentium/core/voice. The recognition adapter is marked preview; validate the selected language, endpoint, and account configuration before deployment.

Configure recognition and synthesis

The host supplies a speaker valid for the selected synthesis model. Language belongs in the session/pipeline configuration rather than the adapter constructor.

Language and turn handling

SARVAM_LANGUAGES is the recognition allowlist, including auto, hi-IN, and en-IN. Synthesis supports a narrower list: en-IN, hi-IN, bn-IN, kn-IN, ml-IN, mr-IN, od-IN, pa-IN, ta-IN, te-IN, and gu-IN. Recognition uses or-IN while synthesis uses od-IN; do not assume the two lists are interchangeable. The recognizer uses manual endpointing. Send ordered audio frames, consume partial/final transcripts, and flush at the host-detected turn boundary. Synthesis accepts streaming text, then flushes the generation. Close both sessions on completion or interruption. Choose compatible input and output rates in the streaming pipeline. The host must perform any required resampling or codec conversion, especially when connecting an 8 kHz phone stream to another provider.