> ## Documentation Index
> Fetch the complete documentation index at: https://docs.agentium.in/llms.txt
> Use this file to discover all available pages before exploring further.

# Sarvam speech adapters

> Configure recognition modes, language codes, and Bulbul synthesis for a streaming pipeline.

Import `SarvamRecognizer` and `SarvamSynthesizer` from `@agentium/core/voice`. The recognition adapter is marked **preview**; validate the selected language, endpoint, and account configuration before deployment.

```bash theme={null}
npm install @agentium/core@4.0.0 ws
export SARVAM_API_KEY="your-key"
```

## Configure recognition and synthesis

```typescript theme={null}
import { SarvamRecognizer, SarvamSynthesizer, SARVAM_LANGUAGES } from "@agentium/core/voice";

export function createSpeech(speaker: string) {
  return {
    recognizer: new SarvamRecognizer({ model: "saaras:v4", mode: "transcribe" }),
    synthesizer: new SarvamSynthesizer({ model: "bulbul:v3", speaker }),
  };
}
console.log("Recognition language codes:", SARVAM_LANGUAGES);
```

The host supplies a speaker valid for the selected synthesis model. Language belongs in the session/pipeline configuration rather than the adapter constructor.

| Setting | Supported by v4 |
| - | - |
| Recognition model | `saaras:v3-realtime` (default), `saaras:v4` |
| Recognition mode | `transcribe`, `translate`, `verbatim`, `translit`, `codemix` |
| Recognition audio | Mono PCM16, μ-law, or A-law at 8 or 16 kHz |
| Synthesis model | `bulbul:v2`, `bulbul:v3` (default) |
| Synthesis audio | Mono PCM16 at 8, 16, 22.05, or 24 kHz |

## Language and turn handling

`SARVAM_LANGUAGES` is the recognition allowlist, including `auto`, `hi-IN`, and `en-IN`. Synthesis supports a narrower list: `en-IN`, `hi-IN`, `bn-IN`, `kn-IN`, `ml-IN`, `mr-IN`, `od-IN`, `pa-IN`, `ta-IN`, `te-IN`, and `gu-IN`. Recognition uses `or-IN` while synthesis uses `od-IN`; do not assume the two lists are interchangeable.

The recognizer uses manual endpointing. Send ordered audio frames, consume partial/final transcripts, and flush at the host-detected turn boundary. Synthesis accepts streaming text, then flushes the generation. Close both sessions on completion or interruption.

Choose compatible input and output rates in the [streaming pipeline](/voice/streaming). The host must perform any required resampling or codec conversion, especially when connecting an 8 kHz phone stream to another provider.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.