# GPU batching

Source: https://docs.vieneu.io/docs/sdk/fast-mode

On a CUDA GPU v3 Turbo batches automatically. A long `infer()` batches its own chunks; `infer_batch()` batches many texts in one forward pass. Same API, no code change.

:::note Renamed page
This page used to describe the LMDeploy "fast mode" for VieNeu-TTS v2. That backend is legacy now; see the bottom of the page.
:::

## infer_batch

```python
from vieneu import Vieneu

tts = Vieneu()                                   # CUDA → PyTorch engine
texts = [
    "Chào cả nhà, hôm nay mình sẽ hướng dẫn các bạn cách cài đặt bộ giọng đọc mới.",
    "Giọng nghe cực kỳ tự nhiên và truyền cảm.",
    "Nếu thấy hữu ích, nhớ để lại một lượt thích nhé!",
] * 10

audios = tts.infer_batch(texts, voice="Minh Quân Pro")
for i, a in enumerate(audios):
    tts.save(a, f"batch_{i}.wav")
```

- Chunks from every text share each forward step, which is where the throughput win comes from.
- Batch size caps at `max_batch_size` (default 32). Tune with `Vieneu(max_batch_size=64)` or `infer_batch(..., batch_size=64)`; `batch_size=1` disables batching.
- On CPU `infer_batch` still works, just sequentially, so there is no gain.
- For real-time playback use [`infer_stream`](/docs/sdk/streaming) instead. It is the streaming twin and serves 16 concurrent listeners on one GPU.

## CUDA graphs (3.7.0+)

Every audio frame is a single CUDA graph replay: acoustic decoder, sampling, repetition penalty and the backbone step in one shot. No `torch.compile`, no C++ toolchain.

Measured on an RTX 3060:

| Input | Audio | Wall time | Before 3.7.0 |
|---|---|---|---|
| One sentence | 3.5 s | 0.36 s | 2.3 s |
| Two-chunk paragraph | 19 s | 1.4 s | 8.7 s |
| 16 chunks | 154 s | 2.8 s (RTF 0.02) | 16.7 s |

The first call for each batch size pays ~0.5 s to capture the graph, then keeps it. Servers can call `tts.warm_fused()` at start-up. `VIENEU_FUSED_FRAME=0` restores the plain loop if you need to debug.

## Requirements

- NVIDIA GPU. An RTX 3060 (12 GB) gives the numbers above; ~6 GB is enough for inference.
- Install per [Install & backends](/docs/sdk/standard-mode#gpu-cuda): CUDA torch 2.8.0 first on Windows, `transformers==4.57.6`, then `vieneu`.

## Fine-tuned models

A LoRA-merged v3 Turbo keeps the full API, batching included:

```python
tts = Vieneu(mode="v3turbo", backbone_repo="finetune/output/my_voice/merged")
```

See [Fine-tuning](/docs/advanced/fine-tuning) and the repo's [`finetune/README.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/finetune/README.md).

## Legacy: LMDeploy "fast" mode (v2 only)

`Vieneu(mode="fast")` loads VieNeu-TTS v1/v2 through LMDeploy. It is kept for existing deployments (`pip install "vieneu[legacy]"`, or `uv sync --group gpu` in the repo) and receives no updates. v3 Turbo on PyTorch is faster and needs no extra runtime.
