LiveKitVoiceTransport implements the media side of StreamingVoicePipeline. It is independent of the LiveKit SIP call adapter.
transport in StreamingVoicePipeline. Feed received room frames through fromLiveKitAudioFrame with a monotonic sequence, turn ID, and generation ID before passing them to the pipeline.
Format and backpressure
The source and input must use mono PCM16. Output rate comes from the source; input rate comes frominputFormat. Resample explicitly when these differ from the selected speech adapters. fromLiveKitAudioFrame copies input samples rather than retaining native frame storage.
queuedDuration is measured in milliseconds. Capture is serialized and bounded by queue time, pending frame count, and timeout. A stuck capture poisons that transport; create a replacement after handling the host source rather than overlapping capture calls.
Playback and ownership
The transport reportsplaybackAcknowledgements: false. Offering a frame to the source does not prove the caller heard it. Without host playback evidence, history remains delivery-unknown.
Clear a cancelled generation to discard queued audio. Close the pipeline and transport when done. With the default ownsSource: false, your host separately unpublishes the track, closes the source, disconnects the room, and disposes native SDK resources. Opt into source ownership only when the transport truly owns it.