Install & backends
One package, two engines. Vieneu() picks the engine from your hardware; the API is the same on both.
| You have | Engine | Install | Notes |
|---|---|---|---|
| CPU only, macOS | ONNX Runtime, torch-free | pip install vieneu | 48 kHz v3 Turbo, cloning and emotion cues included. PyTorch is never installed. |
| NVIDIA GPU | PyTorch (CUDA) | pip install "vieneu[cuda]" | Batched automatically; every frame is one CUDA graph since 3.7.0. |
| Weak CPU / ARM board | ONNX, v3 Nano | pip install vieneu + Vieneu(mode="v3nano") | ~3× faster than Turbo fp32, 24 kHz, lower quality (esp. English). |
CPU (default)
pip install vieneu
from vieneu import Vieneu
tts = Vieneu() # v3 Turbo, ONNX, fp32
audio = tts.infer("Xin chào bạn", voice="Minh Quân Pro")
tts.save(audio, "output.wav") # 48 kHz WAV
- Precision.
fp32is the default for maximum fidelity.Vieneu(precision="int8")is ~1.6× faster and ~4× smaller, but needs a CPU with VNNI (AVX-512 VNNI / AVX-VNNI). On older CPUs int8 can produce garbled audio; if that happens, go back to fp32 or try v3 Nano. - Fastest CPU install. From a checkout of the repo,
uv syncreproduces the locked environment that pins the optimised ONNX Runtime build. It is measurably faster than a plainpip install. - Apple Silicon. Use the CPU/ONNX path. It is faster than the MPS/PyTorch build for v3 Turbo.
GPU (CUDA)
On Linux the PyPI torch wheel already ships CUDA:
pip install "vieneu[cuda]"
On Windows install the CUDA torch first, then a pinned transformers, then vieneu:
pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
pip install "transformers==4.57.6"
pip install vieneu
tts = Vieneu() # CUDA detected → PyTorch engine
tts = Vieneu(backend="onnx") # force the CPU engine on a GPU machine
precision only applies to the ONNX path and is ignored on GPU. Throughput numbers and batching are on the GPU batching page.
v3 Nano (preview)
A 48M-parameter flow-matching model for hardware where Turbo is too slow. Same cloning API, 11 preset voices, 24 kHz output, no frame-level streaming (chunks arrive whole).
tts = Vieneu(mode="v3nano")
audio = tts.infer("Bản nhanh cho máy rất yếu.", voice="Ái Hân", steps=8, sway=-1)
Measured on a 12th-gen i7 desktop, 6 ONNX threads (RTF = compute ÷ audio duration, lower is faster):
| Engine | RTF | Sample rate | Load |
|---|---|---|---|
| v3 Turbo fp32 (CPU default) | 0.62 | 48 kHz | ~19 s |
| v3 Turbo int8 | 0.37 | 48 kHz | ~14 s |
| v3 Nano, 16 steps (default) | 0.22 | 24 kHz | ~3 s |
| v3 Nano, 8 steps, sway −1 | 0.11 | 24 kHz | ~3 s |
Knobs: steps (16 default, 8 ≈ 2× faster), cfg (3.0 default, 0 halves compute but hurts intelligibility), speed, seed, threads.
Models and cache
Backbone and codec weights download from Hugging Face on first use and are cached under ~/.cache/huggingface/hub/. Turbo's cloning pipeline runs on onnxruntime + soxr + kaldi-native-fbank; Nano fetches its three cloning graphs (~110 MB) the first time you clone.
Legacy backends (v1 / v2)
The GGUF (llama-cpp) and LMDeploy backends serve VieNeu-TTS v1/v2 only and are no longer updated. They live behind uv sync --group gpu in the repo and pip install "vieneu[legacy]". New projects should stay on v3 Turbo.