Streaming
v3 Turbo streams frame by frame on both engines. Iterate infer_stream and play or write each chunk as it arrives.
from vieneu import Vieneu
tts = Vieneu() # GPU → PyTorch + stream scheduler; CPU → ONNX
for chunk in tts.infer_stream("Xin chào các bạn!", voice="Mai Anh"):
play(chunk) # np.float32 @ 48 kHz
Latency and concurrency
| Engine | First audio | Concurrent streams |
|---|---|---|
| GPU (PyTorch), RTX 3060 | ~115 ms | 16 (32 max), each at RTF ≈ 0.5–0.6 |
| CPU (ONNX) int8 | ~140 ms | 2 |
| CPU (ONNX) fp32 | ~300 ms | 1 |
On a GPU the streams share one CUDA graph through continuous batching. Calling infer_stream from many threads at once is the intended way to serve many listeners; Vieneu(max_streams=16) sets the ceiling. The first request after the GPU has idled pays an extra 100–300 ms until it clocks up.
Every measurement (TTFA and RTF against max_streams, estimates for smaller GPUs, CPU numbers) is in the repo's docs/streaming.md.
Chunk format
Each chunk is a numpy.float32 array at 48 kHz, mono. Convert to 16-bit PCM for most audio devices:
import numpy as np
pcm16 = (np.clip(chunk, -1, 1) * 32767).astype(np.int16).tobytes()
Serving over HTTP
To expose streaming to other processes or languages, run the OpenAI-compatible server from the repo. It streams pcm/wav as chunked body or SSE from POST /v1/audio/speech, so the OpenAI SDK, Pipecat and LiveKit work by changing base_url. See OpenAI-compatible server.
v3 Nano
Nano has no frame-level streaming: infer_stream yields one finished chunk at a time. Use Turbo when time-to-first-audio matters.
Hosted alternative
The same streaming shape is available without a GPU from the Cloud API streaming endpoint.