Skip to main content

Streaming

v3 Turbo streams frame by frame on both engines. Iterate infer_stream and play or write each chunk as it arrives.

from vieneu import Vieneu

tts = Vieneu() # GPU → PyTorch + stream scheduler; CPU → ONNX
for chunk in tts.infer_stream("Xin chào các bạn!", voice="Mai Anh"):
play(chunk) # np.float32 @ 48 kHz

Latency and concurrency​

EngineFirst audioConcurrent streams
GPU (PyTorch), RTX 3060~115 ms16 (32 max), each at RTF ≈ 0.5–0.6
CPU (ONNX) int8~140 ms2
CPU (ONNX) fp32~300 ms1

On a GPU the streams share one CUDA graph through continuous batching. Calling infer_stream from many threads at once is the intended way to serve many listeners; Vieneu(max_streams=16) sets the ceiling. The first request after the GPU has idled pays an extra 100–300 ms until it clocks up.

Every measurement (TTFA and RTF against max_streams, estimates for smaller GPUs, CPU numbers) is in the repo's docs/streaming.md.

Chunk format​

Each chunk is a numpy.float32 array at 48 kHz, mono. Convert to 16-bit PCM for most audio devices:

import numpy as np

pcm16 = (np.clip(chunk, -1, 1) * 32767).astype(np.int16).tobytes()

Serving over HTTP​

To expose streaming to other processes or languages, run the OpenAI-compatible server from the repo. It streams pcm/wav as chunked body or SSE from POST /v1/audio/speech, so the OpenAI SDK, Pipecat and LiveKit work by changing base_url. See OpenAI-compatible server.

v3 Nano​

Nano has no frame-level streaming: infer_stream yields one finished chunk at a time. Use Turbo when time-to-first-audio matters.

Hosted alternative​

The same streaming shape is available without a GPU from the Cloud API streaming endpoint.