Skip to main content

GPU batching

On a CUDA GPU v3 Turbo batches automatically. A long infer() batches its own chunks; infer_batch() batches many texts in one forward pass. Same API, no code change.

Renamed page

This page used to describe the LMDeploy "fast mode" for VieNeu-TTS v2. That backend is legacy now; see the bottom of the page.

infer_batch​

from vieneu import Vieneu

tts = Vieneu() # CUDA → PyTorch engine
texts = [
"Chào cả nhà, hôm nay mình sẽ hướng dẫn các bạn cách cài đặt bộ giọng đọc mới.",
"Giọng nghe cực kỳ tự nhiên và truyền cảm.",
"Nếu thấy hữu ích, nhớ để lại một lượt thích nhé!",
] * 10

audios = tts.infer_batch(texts, voice="Minh Quân Pro")
for i, a in enumerate(audios):
tts.save(a, f"batch_{i}.wav")
  • Chunks from every text share each forward step, which is where the throughput win comes from.
  • Batch size caps at max_batch_size (default 32). Tune with Vieneu(max_batch_size=64) or infer_batch(..., batch_size=64); batch_size=1 disables batching.
  • On CPU infer_batch still works, just sequentially, so there is no gain.
  • For real-time playback use infer_stream instead. It is the streaming twin and serves 16 concurrent listeners on one GPU.

CUDA graphs (3.7.0+)​

Every audio frame is a single CUDA graph replay: acoustic decoder, sampling, repetition penalty and the backbone step in one shot. No torch.compile, no C++ toolchain.

Measured on an RTX 3060:

InputAudioWall timeBefore 3.7.0
One sentence3.5 s0.36 s2.3 s
Two-chunk paragraph19 s1.4 s8.7 s
16 chunks154 s2.8 s (RTF 0.02)16.7 s

The first call for each batch size pays ~0.5 s to capture the graph, then keeps it. Servers can call tts.warm_fused() at start-up. VIENEU_FUSED_FRAME=0 restores the plain loop if you need to debug.

Requirements​

  • NVIDIA GPU. An RTX 3060 (12 GB) gives the numbers above; ~6 GB is enough for inference.
  • Install per Install & backends: CUDA torch 2.8.0 first on Windows, transformers==4.57.6, then vieneu.

Fine-tuned models​

A LoRA-merged v3 Turbo keeps the full API, batching included:

tts = Vieneu(mode="v3turbo", backbone_repo="finetune/output/my_voice/merged")

See Fine-tuning and the repo's finetune/README.md.

Legacy: LMDeploy "fast" mode (v2 only)​

Vieneu(mode="fast") loads VieNeu-TTS v1/v2 through LMDeploy. It is kept for existing deployments (pip install "vieneu[legacy]", or uv sync --group gpu in the repo) and receives no updates. v3 Turbo on PyTorch is faster and needs no extra runtime.