Introduction
VieNeu turns Vietnamese text into natural speech.
Add https://api.vieneu.io/mcp to your AI assistant, sign in with your VieNeu
account, and ask it to read in a Vietnamese voice — the audio plays right in the
chat. Cursor and VS Code install it in one click.
Set it up in a minute →
To build VieNeu into your own code, there are two ways to use it, and they are different products — start by picking one.
Which one do you want?
☁️ Cloud API — api.vieneu.io
Send text over HTTPS, get audio back. Nothing to install, no GPU, no model download. Billed per character.
This is what you want if you are adding a Vietnamese voice to a product: an
app, a website, a voice agent, a dubbing pipeline. There is an OpenAI-compatible
endpoint, so if your code already calls OpenAI's /v1/audio/speech, pointing it
at VieNeu is a base URL and an API key.
- Quickstart — a key, a voice, one call, audio out
- Cloud API overview — auth, billing, every endpoint
- OpenAI-compatible endpoint — the fastest way in
💻 On-device SDK — the Python package
Runs the model on your own machine. No network call per request, no per-character cost, and the text never leaves your hardware — in exchange you supply the hardware and the setup. Documented in the SDK and Getting Started sections of this site, and summarized below.
The Cloud API and the SDK share a name and a voice catalogue. They do not share an interface: different install, different authentication, different request shapes, different billing. Code written against one does not run against the other, and neither section's documentation applies to the other. Pick the one you are actually using and stay in it.
The on-device SDK
VieNeu-TTS is an advanced on-device Vietnamese Text-to-Speech (TTS) system with instant voice cloning.
Give it text, it speaks it back in natural Vietnamese — fully offline, no cloud API needed.
Key Features
- Instant Voice Cloning — Clone any voice with just 3-5 seconds of reference audio
- Code-switching — Seamless transitions between Vietnamese and English
- Real-time Streaming — Start audio playback before the entire sentence is finished
- Multiple Backends — PyTorch (GPU), GGUF quantized (CPU), LMDeploy (fast GPU), Remote API
- Production Ready — 24 kHz waveform generation, audio watermarking
How It Works
VieNeu-TTS uses a causal language model to generate speech. The core pipeline:
Text → Normalize → Phonemize (eSpeak NG) → LLM generates speech tokens → Codec decodes to audio
- Text normalization — Converts numbers, abbreviations, punctuation to spoken form
- Phonemization — eSpeak NG converts text to pronunciation symbols
- Token generation — A transformer LLM predicts discrete speech tokens
- Audio decoding — NeuCodec converts tokens into a 24kHz waveform
Models
| Model | Format | Quality | Speed |
|---|---|---|---|
| VieNeu-TTS (0.5B) | PyTorch | Best | Very Fast (GPU) |
| VieNeu-TTS-0.3B | PyTorch | Great | Ultra Fast (2x) |
| GGUF Q8 variants | GGUF | Great | Fast (CPU) |
| GGUF Q4 variants | GGUF | Good | Very Fast (CPU) |
All models are hosted on HuggingFace and auto-downloaded on first use.
Quick Start
git clone https://github.com/pnnbao97/VieNeu-TTS.git
cd VieNeu-TTS
uv sync
uv run vieneu-web
Open http://127.0.0.1:7860 and start generating speech.