Skip to main content

Introduction

VieNeu turns Vietnamese text into natural speech.

Using Claude, ChatGPT or Cursor? No code needed

Add https://api.vieneu.io/mcp to your AI assistant, sign in with your VieNeu account, and ask it to read in a Vietnamese voice — the audio plays right in the chat. Cursor and VS Code install it in one click. Set it up in a minute →

To build VieNeu into your own code, there are two ways to use it, and they are different products — start by picking one.

Which one do you want?​

☁️ Cloud API — api.vieneu.io​

Send text over HTTPS, get audio back. Nothing to install, no GPU, no model download. Billed per character.

This is what you want if you are adding a Vietnamese voice to a product: an app, a website, a voice agent, a dubbing pipeline. There is an OpenAI-compatible endpoint, so if your code already calls OpenAI's /v1/audio/speech, pointing it at VieNeu is a base URL and an API key.

💻 On-device SDK — the Python package​

Runs the model on your own machine. No network call per request, no per-character cost, and the text never leaves your hardware — in exchange you supply the hardware and the setup. Documented in the SDK and Getting Started sections of this site, and summarized below.

They are separate products

The Cloud API and the SDK share a name and a voice catalogue. They do not share an interface: different install, different authentication, different request shapes, different billing. Code written against one does not run against the other, and neither section's documentation applies to the other. Pick the one you are actually using and stay in it.


The on-device SDK​

VieNeu-TTS is an advanced on-device Vietnamese Text-to-Speech (TTS) system with instant voice cloning.

Give it text, it speaks it back in natural Vietnamese — fully offline, no cloud API needed.

Key Features​

  • Instant Voice Cloning — Clone any voice with just 3-5 seconds of reference audio
  • Code-switching — Seamless transitions between Vietnamese and English
  • Real-time Streaming — Start audio playback before the entire sentence is finished
  • Multiple Backends — PyTorch (GPU), GGUF quantized (CPU), LMDeploy (fast GPU), Remote API
  • Production Ready — 24 kHz waveform generation, audio watermarking

How It Works​

VieNeu-TTS uses a causal language model to generate speech. The core pipeline:

Text → Normalize → Phonemize (eSpeak NG) → LLM generates speech tokens → Codec decodes to audio
  1. Text normalization — Converts numbers, abbreviations, punctuation to spoken form
  2. Phonemization — eSpeak NG converts text to pronunciation symbols
  3. Token generation — A transformer LLM predicts discrete speech tokens
  4. Audio decoding — NeuCodec converts tokens into a 24kHz waveform

Models​

ModelFormatQualitySpeed
VieNeu-TTS (0.5B)PyTorchBestVery Fast (GPU)
VieNeu-TTS-0.3BPyTorchGreatUltra Fast (2x)
GGUF Q8 variantsGGUFGreatFast (CPU)
GGUF Q4 variantsGGUFGoodVery Fast (CPU)

All models are hosted on HuggingFace and auto-downloaded on first use.

Quick Start​

git clone https://github.com/pnnbao97/VieNeu-TTS.git
cd VieNeu-TTS
uv sync
uv run vieneu-web

Open http://127.0.0.1:7860 and start generating speech.