Skip to main content

Install & backends

One package, two engines. Vieneu() picks the engine from your hardware; the API is the same on both.

You haveEngineInstallNotes
CPU only, macOSONNX Runtime, torch-freepip install vieneu48 kHz v3 Turbo, cloning and emotion cues included. PyTorch is never installed.
NVIDIA GPUPyTorch (CUDA)pip install "vieneu[cuda]"Batched automatically; every frame is one CUDA graph since 3.7.0.
Weak CPU / ARM boardONNX, v3 Nanopip install vieneu + Vieneu(mode="v3nano")~3× faster than Turbo fp32, 24 kHz, lower quality (esp. English).

CPU (default)​

pip install vieneu
from vieneu import Vieneu

tts = Vieneu() # v3 Turbo, ONNX, fp32
audio = tts.infer("Xin chào bạn", voice="Minh Quân Pro")
tts.save(audio, "output.wav") # 48 kHz WAV
  • Precision. fp32 is the default for maximum fidelity. Vieneu(precision="int8") is ~1.6× faster and ~4× smaller, but needs a CPU with VNNI (AVX-512 VNNI / AVX-VNNI). On older CPUs int8 can produce garbled audio; if that happens, go back to fp32 or try v3 Nano.
  • Fastest CPU install. From a checkout of the repo, uv sync reproduces the locked environment that pins the optimised ONNX Runtime build. It is measurably faster than a plain pip install.
  • Apple Silicon. Use the CPU/ONNX path. It is faster than the MPS/PyTorch build for v3 Turbo.

GPU (CUDA)​

On Linux the PyPI torch wheel already ships CUDA:

pip install "vieneu[cuda]"

On Windows install the CUDA torch first, then a pinned transformers, then vieneu:

pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
pip install "transformers==4.57.6"
pip install vieneu
tts = Vieneu()                       # CUDA detected → PyTorch engine
tts = Vieneu(backend="onnx") # force the CPU engine on a GPU machine

precision only applies to the ONNX path and is ignored on GPU. Throughput numbers and batching are on the GPU batching page.

v3 Nano (preview)​

A 48M-parameter flow-matching model for hardware where Turbo is too slow. Same cloning API, 11 preset voices, 24 kHz output, no frame-level streaming (chunks arrive whole).

tts = Vieneu(mode="v3nano")
audio = tts.infer("Bản nhanh cho máy rất yếu.", voice="Ái Hân", steps=8, sway=-1)

Measured on a 12th-gen i7 desktop, 6 ONNX threads (RTF = compute ÷ audio duration, lower is faster):

EngineRTFSample rateLoad
v3 Turbo fp32 (CPU default)0.6248 kHz~19 s
v3 Turbo int80.3748 kHz~14 s
v3 Nano, 16 steps (default)0.2224 kHz~3 s
v3 Nano, 8 steps, sway −10.1124 kHz~3 s

Knobs: steps (16 default, 8 ≈ 2× faster), cfg (3.0 default, 0 halves compute but hurts intelligibility), speed, seed, threads.

Models and cache​

Backbone and codec weights download from Hugging Face on first use and are cached under ~/.cache/huggingface/hub/. Turbo's cloning pipeline runs on onnxruntime + soxr + kaldi-native-fbank; Nano fetches its three cloning graphs (~110 MB) the first time you clone.

Legacy backends (v1 / v2)​

The GGUF (llama-cpp) and LMDeploy backends serve VieNeu-TTS v1/v2 only and are no longer updated. They live behind uv sync --group gpu in the repo and pip install "vieneu[legacy]". New projects should stay on v3 Turbo.