> ## Documentation Index
> Fetch the complete documentation index at: https://docs.agentium.in/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI voice adapters

> Choose native realtime conversation or dedicated streaming transcription.

Agentium exposes two different OpenAI voice adapters from `@agentium/core/voice`: `OpenAIRealtimeProvider` for a native voice conversation, and `OpenAIStreamingRecognizer` for transcription in a composable pipeline.

```bash theme={null}
npm install @agentium/core@4.0.0 ws
export OPENAI_API_KEY="your-key"
```

## Native realtime

```typescript theme={null}
import { OpenAIRealtimeProvider, VoiceAgent } from "@agentium/core/voice";

export function createVoiceAgent() {
  return new VoiceAgent({
    name: "support-voice",
    provider: new OpenAIRealtimeProvider(),
    instructions: "Give short spoken answers. Ask one question at a time.",
    turnDetection: { type: "semantic_vad", eagerness: "low" },
    reasoningEffort: "low",
  });
}
```

Construction does not connect. Call `voice.connect({ tenantId, userId, sessionId })`, attach audio/transcript callbacks, send negotiated audio, and close the returned session when the conversation ends. The [voice session guide](/voice/overview) shows the connection lifecycle.

V4 defaults to `gpt-realtime-2.1`, with `gpt-transcribe` input transcription. It uses nested GA audio configuration, accepts supported PCM/G.711 formats, and rejects unsupported sampling parameters and remote reusable prompts. Do not substitute a text-only model ID.

## Streaming recognition

```typescript theme={null}
import { OpenAIStreamingRecognizer } from "@agentium/core/voice";

export const recognizer = new OpenAIStreamingRecognizer({
  model: "gpt-live-transcribe",
  delay: "low",
  readyTimeoutMs: 10_000,
  turnTimeoutMs: 30_000,
  maxPendingTurns: 8,
});
```

This recognizer accepts mono PCM16 at **24 kHz** and emits partial/final `TranscriptEvent` values. It accepts `gpt-live-transcribe` and `gpt-transcribe`; `delay` is valid only for the live model. Credentials are resolved when opening a session.

Use `open(config, signal)`, consume `session.events` concurrently, send monotonically sequenced audio frames, and call `flush()` at a host-detected turn boundary. Always close the recognizer session. The readiness and committed-turn deadlines bound waiting; they are not audio playback deadlines.

Pass this recognizer to [StreamingVoicePipeline](/voice/streaming) with a compatible transport, brain, and synthesizer. For browser credential minting and SDP exchange, use the separately authenticated [voice gateway](/voice/gateway).

OpenAI recovery creates a fresh provider conversation only when explicitly allowed. It does not replay audio or tool results; see [recovery](/voice/recovery).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.