GPU batching
On a CUDA GPU v3 Turbo batches automatically. A long infer() batches its own chunks; infer_batch() batches many texts in one forward pass. Same API, no code change.
This page used to describe the LMDeploy "fast mode" for VieNeu-TTS v2. That backend is legacy now; see the bottom of the page.
infer_batch
from vieneu import Vieneu
tts = Vieneu() # CUDA → PyTorch engine
texts = [
"Chào cả nhà, hôm nay mình sẽ hướng dẫn các bạn cách cài đặt bộ giọng đọc mới.",
"Giọng nghe cực kỳ tự nhiên và truyền cảm.",
"Nếu thấy hữu ích, nhớ để lại một lượt thích nhé!",
] * 10
audios = tts.infer_batch(texts, voice="Minh Quân Pro")
for i, a in enumerate(audios):
tts.save(a, f"batch_{i}.wav")
- Chunks from every text share each forward step, which is where the throughput win comes from.
- Batch size caps at
max_batch_size(default 32). Tune withVieneu(max_batch_size=64)orinfer_batch(..., batch_size=64);batch_size=1disables batching. - On CPU
infer_batchstill works, just sequentially, so there is no gain. - For real-time playback use
infer_streaminstead. It is the streaming twin and serves 16 concurrent listeners on one GPU.
CUDA graphs (3.7.0+)
Every audio frame is a single CUDA graph replay: acoustic decoder, sampling, repetition penalty and the backbone step in one shot. No torch.compile, no C++ toolchain.
Measured on an RTX 3060:
| Input | Audio | Wall time | Before 3.7.0 |
|---|---|---|---|
| One sentence | 3.5 s | 0.36 s | 2.3 s |
| Two-chunk paragraph | 19 s | 1.4 s | 8.7 s |
| 16 chunks | 154 s | 2.8 s (RTF 0.02) | 16.7 s |
The first call for each batch size pays ~0.5 s to capture the graph, then keeps it. Servers can call tts.warm_fused() at start-up. VIENEU_FUSED_FRAME=0 restores the plain loop if you need to debug.
Requirements
- NVIDIA GPU. An RTX 3060 (12 GB) gives the numbers above; ~6 GB is enough for inference.
- Install per Install & backends: CUDA torch 2.8.0 first on Windows,
transformers==4.57.6, thenvieneu.
Fine-tuned models
A LoRA-merged v3 Turbo keeps the full API, batching included:
tts = Vieneu(mode="v3turbo", backbone_repo="finetune/output/my_voice/merged")
See Fine-tuning and the repo's finetune/README.md.
Legacy: LMDeploy "fast" mode (v2 only)
Vieneu(mode="fast") loads VieNeu-TTS v1/v2 through LMDeploy. It is kept for existing deployments (pip install "vieneu[legacy]", or uv sync --group gpu in the repo) and receives no updates. v3 Turbo on PyTorch is faster and needs no extra runtime.