# Tổng quan SDK

Source: https://docs.vieneu.io/vi/docs/sdk/overview

Gói `vieneu` chạy **VieNeu-TTS v3 Turbo** ngay trên máy của bạn. Mặc định dùng engine ONNX không cần torch trên CPU; có GPU CUDA thì tự chuyển sang PyTorch, không phải đổi code.

:::tip Trang này nằm ở đâu trong bức tranh?
SDK là đường **chạy tại máy** — miễn phí, mã nguồn mở (Apache 2.0), dùng phần cứng của bạn. Nếu muốn gọi API có sẵn (kể cả engine **v4** độc quyền với độ giống giọng cao hơn), xem [Cloud API](/docs/cloud-api/overview).
:::

Các trang sau đi sâu từng chủ đề:

- [Cài đặt & backend](/docs/sdk/standard-mode) — `pip install vieneu`, CPU hay GPU, precision, v3 Nano
- [Batch trên GPU](/docs/sdk/fast-mode) — `infer_batch`, CUDA graph, số đo thông lượng
- [Streaming](/docs/sdk/streaming) — `infer_stream`, nhiều stream đồng thời
- [Nhân bản giọng](/docs/sdk/voice-cloning) — `ref_audio`, `add_voice`, `denoise`
- [Server tương thích OpenAI](/docs/sdk/remote-mode) — `/v1/audio/speech` từ repo hoặc Docker, và chế độ `remote` v2 cũ

:::note Phần dưới là tiếng Anh
Nội dung dưới đây được chép tự động từ README của repo mã nguồn mở, mỗi tháng một lần, nên giữ nguyên tiếng Anh để không bao giờ lệch với bản gốc.
:::

<!-- SDK-README:START — generated by scripts/sync-sdk-readme.mjs, do not edit between the markers -->

:::info Source

This section mirrors the **Using the Python SDK** part of the open-source [README](https://github.com/pnnbao97/VieNeu-TTS#readme) and is refreshed automatically (last sync 2026-09-16). If something here disagrees with the README, the README wins — [open an issue](https://github.com/pnnbao97/VieNeu-TTS/issues) there.

:::

The `vieneu` SDK **defaults to VieNeu-TTS v3 Turbo (48 kHz)**. The minimal install is **torch-free**: on CPU everything runs on **ONNX Runtime** (PyTorch is never imported), and on a CUDA machine it auto-switches to the PyTorch engine — where inference is **batched automatically** (same API, no code change).

## Quick Start

**CPU (default)** — torch-free, runs v3 Turbo via ONNX Runtime. Most users want this:
> ⚡**On CPU the backbone runs `fp32` by default** (maximum fidelity). Need more speed? Pass `Vieneu(precision="int8")` — ~1.6× faster and ~4× smaller, but it requires a CPU with VNNI (AVX-512 VNNI / AVX-VNNI); on older CPUs int8 can produce garbled audio. `precision` only affects the CPU/ONNX path; on GPU it's ignored (PyTorch).
>
> 🪶 **Still too slow, or deploying on a phone / ARM board?** Try **[VieNeu-TTS v3 Nano (preview)](#v3-nano)** — `Vieneu(mode="v3nano")`, ~3× faster than Turbo fp32 on CPU (RTF 0.11–0.22 on a desktop CPU), but **noticeably lower quality** (especially English / bilingual), 24 kHz, 11 preset voices + voice cloning. Details and caveats in the [v3 Nano section](#v3-nano) below.

```bash
pip install vieneu
```

**GPU (CUDA)** — only if you have an NVIDIA GPU. On Linux `pip install "vieneu[cuda]"` is enough (PyPI torch ships CUDA there); on Windows install the CUDA torch **first** as below. 
> ℹ️ **How fast is the GPU path?** Since 3.7.0 every audio frame is **one CUDA
> graph** (acoustic decoder + sampling + repetition penalty + backbone step in a
> single replay — no `torch.compile`, no C++ toolchain needed). Measured on an
> RTX 3060: a 3.5 s sentence in **0.36 s**; a 2-chunk paragraph (19 s) in
> **1.4 s**; 16 chunks (154 s) in **2.8 s** (RTF 0.02) — previously 2.3 s /
> 8.7 s / 16.7 s. The first call for each batch size pays ~0.5 s to capture the
> graph (kept afterwards; servers can call `warm_fused()` at start-up).
> `VIENEU_FUSED_FRAME=0` restores the plain loop.

```bash
pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
pip install "transformers==4.57.6"   # pinned — most stable transformers for the GPU SDK
pip install vieneu
```

```python
import time
from vieneu import Vieneu

# Default = v3 Turbo (48 kHz). GPU → PyTorch (auto-detected).
vieneu = Vieneu() # On a GPU machine you can still switch to ONNX/CPU if you prefer: Vieneu(backend="onnx")

# 1. Built-in voice by name — no reference clip needed
print("🔊 Generating speech...")

start_time = time.time()
audio = vieneu.infer("[cười] Trời ơi, cái giọng nó tự nhiên mà nó mượt mà dã man, nghe không khác gì người thật luôn. Giờ thì tha hồ mà quẩy content với cả kho giọng nói đa dạng, đủ mọi sắc thái biểu cảm. Mọi người bật loa lên rồi cùng trải nghiệm thử với mình nhé!", voice="Phạm Tuyên")
elapsed_time = time.time() - start_time

vieneu.save(audio, "output.wav")
print("✅ Saved to output.wav")

# Tính RTF (Real-Time Factor)
sample_rate = 48000
audio_duration = len(audio) / sample_rate
rtf = elapsed_time / audio_duration

print(f"\n⏱️  Thời gian xử lý: {elapsed_time:.3f}s")
print(f"🎵 Thời lượng audio: {audio_duration:.3f}s")
print(f"📊 RTF: {rtf:.4f}  ({'nhanh hơn' if rtf < 1 else 'chậm hơn'} real-time {1/rtf:.2f}x)" if rtf > 0 else "")

# List the built-in voices
voices = vieneu.list_preset_voices()
print(f"\n🎙️  {len(voices)} built-in voices available:")
for label, voice_id in voices:
    print(f"  - {label} ({voice_id})")

# 2. ⚡ Batch on GPU: infer_batch() runs many texts in ONE batched forward — same API.
#    On a CUDA GPU the chunks from every text share each forward step (big throughput
#    win). On CPU it still WORKS (no error) — just sequentially, so there's no batch
#    gain. Batch caps at max_batch_size (default 32; tune via Vieneu(max_batch_size=64)
#    or infer_batch(..., batch_size=64), or batch_size=1 to disable). A single long
#    infer() also auto-batches its own chunks. For real-time use, infer_stream() is the
#    streaming twin (GPU: 16 concurrent streams — see "Streaming" below). Uncomment to
#    try (GPU recommended):
#
# import time
# texts = [
#     "Chào cả nhà, hôm nay mình sẽ hướng dẫn các bạn cách cài đặt và sử dụng bộ giọng đọc mới.",
#     "Giọng nghe cực kỳ tự nhiên và truyền cảm, lại có thể chuyển đổi biểu cảm một cách linh hoạt.",
#     "Nếu thấy hữu ích, các bạn nhớ để lại một lượt thích và chia sẻ video này cho mọi người nhé!",
# ] * 10   # 30 texts — enough to fill the batch and really show the GPU throughput win
# t0 = time.time()
# audios = vieneu.infer_batch(texts, voice="Minh Quân Pro")
# elapsed = time.time() - t0
# total_audio = sum(len(a) for a in audios) / 48_000
# print(f"⚡ {len(texts)} texts | audio {total_audio:.1f}s | wall {elapsed:.1f}s | RTF {elapsed/total_audio:.3f}")
# for i, a in enumerate(audios):
#     vieneu.save(a, f"batch_{i}.wav")
```

### Streaming (real-time) 🔊

> v3 Turbo streams **frame by frame** on both backends. **GPU** (PyTorch): first audio in **~115 ms** and **16 concurrent streams** on one RTX 3060 (continuous batching — one CUDA graph serves every `infer_stream` call, each keeping RTF ≈ 0.5–0.6). **CPU** (ONNX): first audio in ~140 ms (int8) / ~300 ms (fp32), one stream (two with int8). Just iterate `infer_stream`:

```python
from vieneu import Vieneu
vieneu = Vieneu()                                  # GPU → PyTorch + stream scheduler; no GPU → ONNX/CPU
for chunk in vieneu.infer_stream("Xin chào các bạn!", voice="Mai Anh"):
    play(chunk)                                   # np.float32 @ 48 kHz — play/write as it arrives
```

Calling `infer_stream` from many threads at once is the intended way to serve many listeners on a GPU (`Vieneu(max_streams=16)` sets the ceiling).

An **OpenAI-compatible streaming API** (`POST /v1/audio/speech`, `pcm`/`wav`, chunked or SSE — works with the OpenAI SDK, Pipecat, LiveKit, …) is in [`apps/openai_speech.py`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/apps/openai_speech.py):

```bash
# Pick ONE of these — all serve http://localhost:8000/v1/audio/speech
uv run python -m apps.openai_speech                                  # from the repo (auto-detects GPU/CPU)
docker compose -f docker/docker-compose.yml --profile api-gpu up     # or: Docker, GPU
docker compose -f docker/docker-compose.yml --profile api-cpu up     # or: Docker, CPU only
```

📊 **[docs/streaming.md](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/streaming.md)** — every measurement on an RTX 3060 (TTFA / RTF / streams vs `max_streams`), estimates for smaller GPUs, and the CPU numbers. The older browser demo is still at [`apps/web_stream.py`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/apps/web_stream.py).

### Available Voices

The v3 Turbo engine includes **25 preset voices** covering **3 regions** (North, Central, South) with diverse genders and speaking characters. `list_preset_voices()` (and the Web UI / API voice lists) show them in this order:

- ⭐ **Editors' picks** — the 10 we recommend starting with, hand-selected for naturalness and stability: **Adam bựa, Trúc Ly, Anh Khôi, Mai Anh, Minh Quân Pro** *(default; `"Minh Quân"` still works as an alias)*, **Thùy Dung, Thiền Tâm Đức, Ngọc Huyền, Quang Sơn, Ngọc Trân**
- **Northern (Bắc)**: Minh Đức, Phạm Tuyên, Xuân Vĩnh, Thanh Bình, Ngọc Linh, Đoan Trang, Quỳnh Anh, Mạnh Dũng (+ picks above)
- **Central (Trung)**: Quang Sơn, Ngọc Trân
- **Southern (Nam)**: Adam, Thái Sơn, Thục Đoan, Minh Triết, Mỹ Duyên, Đức Trí, Kim Thanh (+ Thùy Dung)

## Reading style — **deprecated** ⚠️

:::warning

**`style` is deprecated on v3 Turbo and has no effect.** The reading style is already
baked into the reference itself (the speaker embedding + reference codes of the preset
voice or of your cloned clip), so every generation follows the reference and comes out
in its natural reading style.

The `style` argument is **still accepted** by `infer`, `infer_stream`, `infer_batch`
and `add_voice` so existing code keeps running — whatever you pass (`"tin_tuc"`,
`"doc_truyen"`, …) is simply ignored. New code should just omit it.

:::

```python
# Old code — still runs, but `style` is ignored
audio = vieneu.infer("Bản tin sáng nay.", voice="Minh Quân Pro", style="tin_tuc")

# New code — pick the reading character through the voice / reference clip instead
audio = vieneu.infer("Bản tin sáng nay.", voice="Minh Quân Pro")
```

## Emotion cues (experimental)

Inline tags are supported anywhere in the text: `[cười]` (chuckle), `[thở dài]` (sigh), `[hắng giọng]` (clear throat).

```python
audio = vieneu.infer("Nghe hay quá đi [cười]. Để mình nói tiếp [hắng giọng].", voice="Minh Quân Pro")
```

## Voice cloning

Clone any voice from a short reference clip. The clip is cleaned up automatically
(background noise removed, and trimmed to ≤ 8 seconds) before cloning — keep
`denoise=True` unless your clip is already clean.

```python
audio = vieneu.infer(
    "Đây là giọng được nhân bản tức thì.",
    ref_audio="my_voice.wav",   # a 3–8s reference clip
    denoise=True,               # default; set False if the clip is already clean
)
vieneu.save(audio, "cloned.wav")
```

## Save & reuse a cloned voice

Register a reference once with `add_voice`, then use it by name like a built-in voice.

```python
# Enroll a voice (denoises + extracts the speaker profile once)
vieneu.add_voice("Giọng của tôi", "my_voice.wav")

# Now reuse it anywhere, including the conversation mode
audio = vieneu.infer("Câu này dùng giọng đã lưu.", voice="Giọng của tôi")

# Persist your voices so they load next session
vieneu.save_voices()                 # writes to the default voices file
# vieneu.remove_voice("Giọng của tôi")

# Add a voice you already cleaned yourself → skip denoising
vieneu.add_voice("Giọng sạch", "already_clean.wav", denoise=False)
```

## Clean up a clip on its own

Get the denoised audio without synthesizing anything (e.g. to inspect or store it):

```python
wav, sr = vieneu.denoise("noisy.wav", out_path="clean.wav")   # 44.1 kHz mono
```

> **Note:** `denoise`, `add_voice`, and voice cloning work on every backend — the
> torch-free CPU/ONNX install included (the whole cloning pipeline runs on
> onnxruntime + soxr + kaldi-native-fbank). **v3 Nano** below clones the same way (its cloning graphs are fetched on first use).

## v3 Nano (preview) — for edge devices / weak CPUs only 🪶 {#v3-nano}
:::warning

**v3 Turbo remains the default and the recommended model.** Use v3 Nano only when Turbo is
too slow on your hardware (old laptops, mini PCs, ARM boards, CPUs without AVX-512/VNNI where
the int8 Turbo build produces garbled audio). Nano is a 48M-parameter flow-matching model
(ONNX, CPU, torch-free) and it **trades quality for speed**:
- **Lower quality than v3 Turbo — most noticeably on English and code-switched (En-Vi) text.**
  Vietnamese is close; English words come out with a Vietnamese accent and are less stable.
- **24 kHz** output (Turbo: 48 kHz).
- **11 preset voices + voice cloning** (`ref_audio`, `add_voice`, `encode_reference` work like Turbo; the three cloning graphs, ~110 MB, download on first use).
- **No frame-level streaming** — `infer_stream` yields one finished chunk at a time.

:::

Measured on the same desktop CPU (12th-gen Intel i7, 6 ONNX Runtime threads, ~9 s of speech):

| Engine | RTF ↓ | Sample rate | Load time |
|---|---|---|---|
| v3 Turbo ONNX fp32 (default on CPU) | 0.62 | 48 kHz | ~19 s |
| v3 Turbo ONNX int8 | 0.37 | 48 kHz | ~14 s |
| **v3 Nano, 16 steps, cfg 3** (default) | **0.22** | 24 kHz | ~3 s |
| **v3 Nano, 8 steps, sway −1** | **0.11** | 24 kHz | ~3 s |

RTF = compute time ÷ audio duration (lower is faster; 0.22 = 4.5× faster than real time). The ratio carries over to slower machines: expect Nano to be roughly **1.7× faster than Turbo int8** and **~3× faster than Turbo fp32**, with a 282 MB download instead of Turbo's.

```python
from vieneu import Vieneu

tts = Vieneu(mode="v3nano")                      # ONNX, CPU, torch-free
audio = tts.infer("Xin chào, mình là giọng đọc của VieNeu Nano.", voice="Minh Quân")
tts.save(audio, "nano.wav")                      # 24 kHz

tts.list_preset_voices()                         # Adam, Ái Hân, Mỹ Duyên, Đức Trí, Hữu Quân, Xuân Tiên, Mai Anh, Trúc Ly, Anh Khôi, Minh Quân, Mạnh Dũng
audio = tts.infer("Bản nhanh cho máy rất yếu.", voice="Ái Hân", steps=8, sway=-1)   # ~2× faster
```

Knobs: `steps` (Euler steps, 16 default; 8 ≈ 2× faster, slightly rougher — pair with `sway=-1`),
`cfg` (classifier-free guidance, 3.0 default; `cfg=0` halves compute but hurts intelligibility),
`speed`, `seed`, `threads`. Emotion cues `[cười]` `[thở dài]` `[hắng giọng]` work as on Turbo.

<!-- SDK-README:END -->
