# Streaming

Source: https://docs.vieneu.io/docs/sdk/streaming

v3 Turbo streams **frame by frame** on both engines. Iterate `infer_stream` and play or write each chunk as it arrives.

```python
from vieneu import Vieneu

tts = Vieneu()                                    # GPU → PyTorch + stream scheduler; CPU → ONNX
for chunk in tts.infer_stream("Xin chào các bạn!", voice="Mai Anh"):
    play(chunk)                                   # np.float32 @ 48 kHz
```

## Latency and concurrency

| Engine | First audio | Concurrent streams |
|---|---|---|
| GPU (PyTorch), RTX 3060 | ~115 ms | 16 (32 max), each at RTF ≈ 0.5–0.6 |
| CPU (ONNX) int8 | ~140 ms | 2 |
| CPU (ONNX) fp32 | ~300 ms | 1 |

On a GPU the streams share one CUDA graph through continuous batching. Calling `infer_stream` from many threads at once is the intended way to serve many listeners; `Vieneu(max_streams=16)` sets the ceiling. The first request after the GPU has idled pays an extra 100–300 ms until it clocks up.

Every measurement (TTFA and RTF against `max_streams`, estimates for smaller GPUs, CPU numbers) is in the repo's [`docs/streaming.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/streaming.md).

## Chunk format

Each chunk is a `numpy.float32` array at 48 kHz, mono. Convert to 16-bit PCM for most audio devices:

```python
import numpy as np

pcm16 = (np.clip(chunk, -1, 1) * 32767).astype(np.int16).tobytes()
```

## Serving over HTTP

To expose streaming to other processes or languages, run the OpenAI-compatible server from the repo. It streams `pcm`/`wav` as chunked body or SSE from `POST /v1/audio/speech`, so the OpenAI SDK, Pipecat and LiveKit work by changing `base_url`. See [OpenAI-compatible server](/docs/sdk/remote-mode).

## v3 Nano

Nano has no frame-level streaming: `infer_stream` yields one finished chunk at a time. Use Turbo when time-to-first-audio matters.

## Hosted alternative

The same streaming shape is available without a GPU from the [Cloud API streaming endpoint](/docs/cloud-api/streaming).
