> ## Documentation Index
> Fetch the complete documentation index at: https://docs.agentium.in/llms.txt
> Use this file to discover all available pages before exploring further.

# Voice gateway

> Connect browser audio to native voice sessions with bounded buffers and playback acknowledgements.

`createVoiceGateway` from `@agentium/transport` connects Socket.IO clients to registered `VoiceAgent` instances. It reserves sessions before asynchronous setup, cancels late connections after disconnect, and bounds queued audio.

## Authenticate connections

When you configure `authMiddleware`, verify credentials and populate `socket.data.auth` with `{ userId, tenantId?, sessionId? }`. Client user/session/API-key overrides are ignored in authenticated mode; missing trusted identity rejects setup. Keep provider credentials on the server.

The text gateway has a separate required [security mode](/transport/socketio). Do not assume configuring one gateway secures the other.

## Client event contract

| Event | Direction | Client behavior |
| - | - | - |
| `voice.started` | Server → client | Read negotiated input/output formats and acknowledgement support |
| `voice.audio` | Server → client | Queue base64 bytes with their `sequence`, `generationId`, and format |
| `voice.playback.ack` | Client → server | Send `{ sequence }` after consuming the frame to release its buffer budget |
| `voice.turn_complete` | Server → client | Generation has finished; playback may still be pending |
| `voice.playback.complete` | Client → server | Send `{ generationId }` only after that generation has actually played |
| `voice.clear` | Server → client | Discard queued playback immediately; do not mark discarded speech as heard |
| `voice.commit` | Client → server | Flush manually committed input |
| `voice.recovery` | Server → client | Pause capture during recovery, then supply new input after recovery |

Sequence acknowledgement and playback confirmation serve different purposes. Only send played-character offsets when backed by audio/text alignment. Send playback completion even if every frame has already been acknowledged.

## Bounds and disconnects

Defaults are **256 KiB per frame**, **1 MiB pending output**, and a **10-second acknowledgement deadline**. Malformed audio, overflow, and stalled acknowledgements terminate the session. Configure Socket.IO server and deployment limits separately.

During recovery, audio and commits are discarded; text receives a retryable error. No queued inputs are replayed after recovery. A disconnect or local abort does not prove an external tool effect stopped.

See [native sessions](/voice/overview) for provider recovery and [streaming voice](/voice/streaming) for a coordinator-owned Agent pipeline.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.