# Install & backends

Source: https://docs.vieneu.io/docs/sdk/standard-mode

One package, two engines. `Vieneu()` picks the engine from your hardware; the API is the same on both.

| You have | Engine | Install | Notes |
|---|---|---|---|
| CPU only, macOS | ONNX Runtime, **torch-free** | `pip install vieneu` | 48 kHz v3 Turbo, cloning and emotion cues included. PyTorch is never installed. |
| NVIDIA GPU | PyTorch (CUDA) | `pip install "vieneu[cuda]"` | Batched automatically; every frame is one CUDA graph since 3.7.0. |
| Weak CPU / ARM board | ONNX, **v3 Nano** | `pip install vieneu` + `Vieneu(mode="v3nano")` | ~3× faster than Turbo fp32, 24 kHz, lower quality (esp. English). |

## CPU (default)

```bash
pip install vieneu
```

```python
from vieneu import Vieneu

tts = Vieneu()                       # v3 Turbo, ONNX, fp32
audio = tts.infer("Xin chào bạn", voice="Minh Quân Pro")
tts.save(audio, "output.wav")        # 48 kHz WAV
```

- **Precision.** `fp32` is the default for maximum fidelity. `Vieneu(precision="int8")` is ~1.6× faster and ~4× smaller, but needs a CPU with VNNI (AVX-512 VNNI / AVX-VNNI). On older CPUs int8 can produce garbled audio; if that happens, go back to fp32 or try v3 Nano.
- **Fastest CPU install.** From a checkout of the repo, `uv sync` reproduces the locked environment that pins the optimised ONNX Runtime build. It is measurably faster than a plain `pip install`.
- **Apple Silicon.** Use the CPU/ONNX path. It is faster than the MPS/PyTorch build for v3 Turbo.

## GPU (CUDA)

On Linux the PyPI torch wheel already ships CUDA:

```bash
pip install "vieneu[cuda]"
```

On Windows install the CUDA torch **first**, then a pinned transformers, then vieneu:

```bash
pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
pip install "transformers==4.57.6"
pip install vieneu
```

```python
tts = Vieneu()                       # CUDA detected → PyTorch engine
tts = Vieneu(backend="onnx")         # force the CPU engine on a GPU machine
```

`precision` only applies to the ONNX path and is ignored on GPU. Throughput numbers and batching are on the [GPU batching](/docs/sdk/fast-mode) page.

## v3 Nano (preview)

A 48M-parameter flow-matching model for hardware where Turbo is too slow. Same cloning API, 11 preset voices, 24 kHz output, no frame-level streaming (chunks arrive whole).

```python
tts = Vieneu(mode="v3nano")
audio = tts.infer("Bản nhanh cho máy rất yếu.", voice="Ái Hân", steps=8, sway=-1)
```

Measured on a 12th-gen i7 desktop, 6 ONNX threads (RTF = compute ÷ audio duration, lower is faster):

| Engine | RTF | Sample rate | Load |
|---|---|---|---|
| v3 Turbo fp32 (CPU default) | 0.62 | 48 kHz | ~19 s |
| v3 Turbo int8 | 0.37 | 48 kHz | ~14 s |
| v3 Nano, 16 steps (default) | 0.22 | 24 kHz | ~3 s |
| v3 Nano, 8 steps, sway −1 | 0.11 | 24 kHz | ~3 s |

Knobs: `steps` (16 default, 8 ≈ 2× faster), `cfg` (3.0 default, 0 halves compute but hurts intelligibility), `speed`, `seed`, `threads`.

## Models and cache

Backbone and codec weights download from Hugging Face on first use and are cached under `~/.cache/huggingface/hub/`. Turbo's cloning pipeline runs on onnxruntime + soxr + kaldi-native-fbank; Nano fetches its three cloning graphs (~110 MB) the first time you clone.

## Legacy backends (v1 / v2)

The GGUF (llama-cpp) and LMDeploy backends serve **VieNeu-TTS v1/v2** only and are no longer updated. They live behind `uv sync --group gpu` in the repo and `pip install "vieneu[legacy]"`. New projects should stay on v3 Turbo.
