# VieNeu documentation — full text
> Vietnamese Text-to-Speech with Instant Voice Cloning
Every documentation page below, in sidebar order, each headed by the URL it
is served from. The API reference is not included here — it is generated
separately at https://docs.vieneu.io/api-reference.md, from
https://docs.vieneu.io/openapi.json.
Index: https://docs.vieneu.io/llms.txt
---
# Introduction
Source: https://docs.vieneu.io/docs/
VieNeu turns Vietnamese text into natural speech.
:::tip Using Claude, ChatGPT or Cursor? No code needed
Add `https://api.vieneu.io/mcp` to your AI assistant, sign in with your VieNeu
account, and ask it to read in a Vietnamese voice — the audio plays right in the
chat. Cursor and VS Code install it in one click.
**[Set it up in a minute →](./integrations/mcp/index.md)**
:::
To build VieNeu into your own code, there are **two ways to use it**, and they
are different products — start by picking one.
## Which one do you want?
### ☁️ Cloud API — `api.vieneu.io`
Send text over HTTPS, get audio back. Nothing to install, no GPU, no model
download. Billed per character.
This is what you want if you are **adding a Vietnamese voice to a product**: an
app, a website, a voice agent, a dubbing pipeline. There is an OpenAI-compatible
endpoint, so if your code already calls OpenAI's `/v1/audio/speech`, pointing it
at VieNeu is a base URL and an API key.
- **[Quickstart](./cloud-api/quickstart.md)** — a key, a voice, one call, audio out
- **[Cloud API overview](./cloud-api/overview.md)** — auth, billing, every endpoint
- **[OpenAI-compatible endpoint](./cloud-api/openai-compatible.md)** — the fastest way in
### 💻 On-device SDK — the Python package
Runs the model on your own machine. No network call per request, no per-character
cost, and the text never leaves your hardware — in exchange you supply the
hardware and the setup. Documented in the **SDK** and **Getting Started** sections
of this site, and summarized below.
:::info They are separate products
The Cloud API and the SDK share a name and a voice catalogue. They do **not**
share an interface: different install, different authentication, different
request shapes, different billing. Code written against one does not run against
the other, and neither section's documentation applies to the other. Pick the one
you are actually using and stay in it.
:::
---
## The on-device SDK
**VieNeu-TTS** is an advanced on-device Vietnamese Text-to-Speech (TTS) system with **instant voice cloning**.
Give it text, it speaks it back in natural Vietnamese — fully offline, no cloud API needed.
## Key Features
- **Instant Voice Cloning** — Clone any voice with just 3-5 seconds of reference audio
- **Code-switching** — Seamless transitions between Vietnamese and English
- **Real-time Streaming** — Start audio playback before the entire sentence is finished
- **Multiple Backends** — PyTorch (GPU), GGUF quantized (CPU), LMDeploy (fast GPU), Remote API
- **Production Ready** — 24 kHz waveform generation, audio watermarking
## How It Works
VieNeu-TTS uses a **causal language model** to generate speech. The core pipeline:
```
Text → Normalize → Phonemize (eSpeak NG) → LLM generates speech tokens → Codec decodes to audio
```
1. **Text normalization** — Converts numbers, abbreviations, punctuation to spoken form
2. **Phonemization** — eSpeak NG converts text to pronunciation symbols
3. **Token generation** — A transformer LLM predicts discrete speech tokens
4. **Audio decoding** — NeuCodec converts tokens into a 24kHz waveform
## Models
| Model | Format | Quality | Speed |
|-------|--------|---------|-------|
| VieNeu-TTS (0.5B) | PyTorch | Best | Very Fast (GPU) |
| VieNeu-TTS-0.3B | PyTorch | Great | Ultra Fast (2x) |
| GGUF Q8 variants | GGUF | Great | Fast (CPU) |
| GGUF Q4 variants | GGUF | Good | Very Fast (CPU) |
All models are hosted on [HuggingFace](https://huggingface.co/pnnbao-ump) and auto-downloaded on first use.
## Quick Start
```bash
git clone https://github.com/pnnbao97/VieNeu-TTS.git
cd VieNeu-TTS
uv sync
uv run vieneu-web
```
Open `http://127.0.0.1:7860` and start generating speech.
---
# MCP server — VieNeu in Claude, ChatGPT and more
Source: https://docs.vieneu.io/docs/integrations/mcp/
VieNeu runs a hosted [Model Context Protocol](https://modelcontextprotocol.io)
(MCP) server. Add one URL to an AI assistant, sign in with your VieNeu account,
and ask for Vietnamese speech in plain language — the assistant finds a voice,
synthesizes the text, and plays it back or hands you a download link.
```
https://api.vieneu.io/mcp
```
Nothing to install. The assistant signs in through your browser (OAuth 2.1) the
first time you use it; tools and apps that cannot sign in can send an
[API key](#using-an-api-key-instead-of-signing-in) instead.
**Before you start:** you need a VieNeu account with an **active token plan**.
Connecting is refused without one, because every synthesis would fail.
## Pick your app
| App | How you connect | Inline player | Guide |
|---|---|---|---|
| **Claude** (web, desktop, mobile) | Add a custom connector, sign in | Yes | [Claude](./claude-ai.md) |
| **ChatGPT** | Developer mode → add an app, sign in | Yes | [ChatGPT](./chatgpt.md) |
| **Claude Code** | One command, then `/mcp` | — | [Claude Code](./claude-code.md) |
| **Cursor** | One click or `mcp.json`; sign in or API key | Yes | [Cursor](./cursor.md) |
| **VS Code** (GitHub Copilot) | One click or `mcp.json`; sign in or API key | Yes (experimental) | [VS Code](./vscode.md) |
| **OpenAI API / Agents SDK** | `mcp` tool in your code | — | [OpenAI API](./openai-api.md) |
| Windsurf (Devin Desktop), Gemini CLI, Codex CLI, Zed, Goose, LM Studio | Config file | Goose only | [Other apps](./other-clients.md) |
Something not working? See [Troubleshooting](./troubleshooting.md).
## What you say to it
> "What can VieNeu do?"
The assistant calls `list_capabilities` and answers with what works here, one
example request for each, and which ones spend tokens. In apps with the
[inline player](#the-inline-player) it is a card: click an example and it is
sent to the chat.
> "Find me a northern female voice for news."
The assistant calls `list_voices`. Free.
> "Read this paragraph with Thu Trang."
The assistant calls `text_to_speech`. In apps with the inline player, a small
player appears right in the chat — press play. There is always a download link
too (valid 24 hours; MP3 for texts up to 1,500 characters). The audio is saved in
your **Library** on vieneu.io. This spends tokens from your plan, exactly like
the same request made through the API.
> "Add a laugh after the first sentence."
The assistant calls `list_emotion_tags` and writes a cue tag such as `[cười]`
into the text before synthesizing.
> "Make an audiobook of this story — a northern woman narrating, a voice for each character."
The assistant scripts your text, chooses the voices, tells you the cost, and
after you agree VieNeu makes one MP3 per chapter in the background. See
[Audiobooks](./audiobooks.md).
## Tools
| Tool | What it does | Cost |
|------|--------------|------|
| `list_capabilities` | What this connection can do, with an example request for each | Free |
| `list_voices` | Search the voices your account can use, including your own cloned voices | Free |
| `text_to_speech` | Synthesize Vietnamese text; returns a download link and the duration | Tokens |
| `get_speech_status` | Check a long job that was still processing; returns the link when done | Free |
| `list_emotion_tags` | Reading styles, and cue tags such as `[cười]` to write into the text | Free |
| `get_token_balance` | Tokens you can spend right now, your plan's caps and when they reset, and roughly how many characters that reads | Free |
Six more tools — `create_audiobook`, `add_audiobook_chapter`, `start_audiobook`,
`get_audiobook`, `list_audiobooks`, `cancel_audiobook` — make audiobooks; they
are described in [Audiobooks](./audiobooks.md#tools).
### `list_voices`
| Argument | Type | Notes |
|---|---|---|
| `search` | string, optional | A few words describing the voice (`nữ miền Bắc trầm ấm`, `male south`), or its name (`Thu Trang`). How the words are read is below. |
| `limit` | 1–100, optional | Default 25. The reply says how many more matched. |
Each line reads `id (gender, region) "Name" - description`. The **id** is what
`text_to_speech` needs, exactly as shown — ids look like names (`Thu Trang`)
and cannot be guessed.
How `search` reads the words:
- **Gender and region** come from the voice's fields, never from its
description: `nữ`, `nam`, `miền Bắc`, `miền Nam`, `miền Trung`, or `female`,
`male`, `north`, `south`, `central`. `nam` on its own means male; write
`miền Nam` for the South.
- **Other words** match whole words of the voice's id, name or description,
ignoring case. A word written with diacritics must match them (`già`, old,
does not find a voice named Gia); written without, it matches either way,
and from 4 letters also as the start of a word.
- **Words that describe no voice**, such as `giọng` or `đọc`, are ignored.
- **A voice's name** (`Thùy An`, `giọng Thu Trang`) puts that voice first.
- **When no voice has every word**, the closest come back instead, those with
the gender and region asked for first, and each line ends with the words
that voice lacks: `[thiếu: trầm, ấm]`.
### `text_to_speech`
| Argument | Type | Notes |
|---|---|---|
| `text` | string | Vietnamese text, up to 50,000 characters. Numbers, dates and English acronyms are handled. |
| `voice` | string, optional | A voice id from `list_voices`. Omitted: the default voice. |
| `speed` | 0.5–2.0, optional | 1.0 is natural pace. |
| `style` | string, optional | A reading style from `list_emotion_tags`. |
Up to 1,500 characters the tool makes one
[`POST /v1/audio/speech`](../../cloud-api/openai-compatible.md) call for **MP3**
and returns its link a moment later. Longer text becomes one
[`POST /v1/tts`](../../cloud-api/overview.md) job (**WAV**) that the tool waits
on for about 50 seconds; a very long text may still be processing then — the
reply gives a job id and tells the assistant to call `get_speech_status` later
instead of synthesizing again (which would be charged again). Either way the
reply carries the link, its format, the duration and the voice — and, once the
audio is ready, how many tokens you have left.
### `get_speech_status`
| Argument | Type | Notes |
|---|---|---|
| `job_id` | string | The id `text_to_speech` returned. |
### `get_token_balance`
No arguments. Ask *"còn bao nhiêu token?"* and the assistant answers with:
- **Usable now** — the most one synthesis can spend at this moment: the
plan's remaining tokens, or less when a daily or weekly cap is lower.
- The plan's remaining / total, each cap and when it resets (Vietnam time),
and when the plan expires.
- The total across all your active plans, if you have more than one — the
next one takes over when this one runs out.
- About how many characters that reads with the default voice engine.
At zero it links to the pricing page. The same numbers are on
[`GET /v1/balance`](../../cloud-api/overview.md#seeing-what-you-were-billed-for)
for scripts.
## The inline player
In apps that support [MCP Apps](https://modelcontextprotocol.io/docs/extensions/apps)
— Claude (web, desktop and mobile), ChatGPT, Cursor, VS Code (experimental) and
Goose — `text_to_speech` and `get_speech_status` results show a player in the
conversation: play/pause, seek, duration, and a download button. While a long
job is still running it shows a small green light and keeps checking
`get_speech_status` by itself (free), switching to the player when the audio is
ready.
Audiobooks get their own player under `start_audiobook` and `get_audiobook`:
the chapters with their progress, played one after another, refreshing itself
while the book is being made.
Claude asks once before showing it — choose **Allow** (or **Always allow**).
Apps without MCP Apps show the text reply with the link instead; nothing else
changes.
## Ready-made prompts {#prompts}
The server also offers five prompts — ready-made requests with a form for
their details. Apps that show MCP prompts list them in a menu: Claude Code as
`/mcp__vieneu__make_audiobook`, VS Code as `/mcp.vieneu.make_audiobook` (the
middle part is the name you gave the server). ChatGPT does not show prompts;
ask *"VieNeu làm được gì?"* there instead.
| Prompt | What it asks for | Details |
|---|---|---|
| `what_can_vieneu_do` | The list of what VieNeu can do | — |
| `read_text` | Read a text aloud | `text`, `voice` (optional, e.g. "nữ miền Bắc") |
| `make_audiobook` | An audiobook from your text, priced before it starts | `text` (optional — paste it after), `narrator` (optional) |
| `find_voice` | Voices matching a description | `description` |
| `check_balance` | Tokens left and the plan's limits | — |
## Using an API key instead of signing in
Scripts, the OpenAI API, and apps that cannot sign in can send one of your
[API keys](../../cloud-api/overview.md#authentication) on every request, in
either header:
```
X-API-Key: vn_sk_...
Authorization: Bearer vn_sk_...
```
Create a key on the **Developer** page of vieneu.io. A config file holding a key
is a password on disk — prefer signing in where the app supports it, and never
commit a key to a repository. Each app's guide shows where the header goes.
## Billing
A tool call costs exactly what the equivalent API call costs: one synthesis per
`text_to_speech` call, charged per character (minimum 50) at the engine's rate,
from your plan's tokens. A failed synthesis is refunded automatically. An
audiobook is billed the same way, chapter by chapter as each one is made, after
you agree to `start_audiobook`. Every other tool is free.
MCP usage appears under **Developer → Usage** like any API traffic, attributed
to the connection (for example "Claude (MCP)").
## See and remove connections
**Developer → Dùng VieNeu trong Claude, ChatGPT** on vieneu.io shows the URL to
copy and lists every app you have connected, with the date and the last use.
**Gỡ kết nối** (Disconnect) signs that app out immediately; it has to ask for
permission again to come back.
## Errors
The tools translate API errors into sentences the assistant can act on.
| What you see | Cause | Fix |
|---|---|---|
| Asked to sign in again | The connection was removed on vieneu.io, or its sign-in expired | Reconnect from the app |
| "Tài khoản không đủ quyền hoặc hết token…" | The plan is out of tokens, expired, or does not include the engine (HTTP 403) | Top up or renew on vieneu.io — retrying does not help |
| "Đang gọi quá nhanh…" | Rate limited (HTTP 429) | Wait the number of seconds quoted |
| "Máy chủ tạo giọng đang bận" | No synthesis worker free (HTTP 503) | Retry in a few seconds |
| "Voice … is not available" | The voice id is not in your catalogue | Pick an id from `list_voices` |
More in [Troubleshooting](./troubleshooting.md).
## What is deliberately not exposed
Voice cloning, dubbing and SRT dubbing are not tools here. They take file
uploads and spend noticeably more per call, and an assistant invoking one by
accident is a bad first experience. Use the web app or the
[Cloud API](../../cloud-api/overview.md) for those.
## For client developers
- Transport: Streamable HTTP, stateless. Protocol revision 2026-07-28, with
2025-era clients served too.
- Authorization: OAuth 2.1 per the MCP authorization spec. An unauthenticated
request gets `401` with
`WWW-Authenticate: Bearer resource_metadata="https://api.vieneu.io/.well-known/oauth-protected-resource"`.
- Metadata: [`/.well-known/oauth-protected-resource`](https://api.vieneu.io/.well-known/oauth-protected-resource)
(RFC 9728) and
[`/.well-known/oauth-authorization-server`](https://api.vieneu.io/.well-known/oauth-authorization-server)
(RFC 8414).
- Clients: Client ID Metadata Documents and Dynamic Client Registration
(`/oauth/register`). PKCE `S256` is always required. A metadata document must
allow `none` (preferred or among `token_endpoint_auth_methods_supported`).
Registration accepts `none`, `client_secret_post` and `client_secret_basic`
and always returns a `client_secret`; it is required only for the two secret
methods. Loopback redirects (`http://127.0.0.1` / `localhost`)
match on any port.
- Scope: `tts` (plus `offline_access` for a refresh token). Access tokens last
one hour; refresh tokens rotate on every use.
- MCP Apps: `ui://vieneu/player.html`, `ui://vieneu/audiobook.html` and
`ui://vieneu/capabilities.html` (`text/html;profile=mcp-app`); media allowed
from the storage origin only. The capabilities card sends an example to the
chat with `ui/message` when the host offers it.
- Prompts: `what_can_vieneu_do`, `read_text`, `make_audiobook`, `find_voice`,
`check_balance`.
---
# Make an audiobook with Claude or ChatGPT
Source: https://docs.vieneu.io/docs/integrations/mcp/audiobooks
Give the assistant a story, a chapter or a whole book — paste it or attach the
file — and ask for an audiobook:
> "Làm sách nói từ truyện này. Giọng dẫn chuyện nữ miền Bắc, mỗi nhân vật một giọng riêng."
The assistant turns your text into a script — the narration, and every line of
dialogue with the character who says it — picks a narrator and a voice for each
character, and tells you what it will cost. Once you agree, VieNeu reads the
book chapter by chapter and masters each chapter into an MP3 ready to publish.
This works in every app on the [MCP server](./index.md) — Claude and ChatGPT show
a player for the chapters right in the chat.
## How it goes
1. **The script.** The assistant creates the book and adds the chapters in
order. It copies your text word for word — its instructions forbid
summarizing or rewriting — and leaves out tables of contents, page numbers
and copyright pages. Free.
2. **The cost.** Before anything is spent it tells you the estimate in tokens:
every character is billed like any synthesis, with the default engine's
rate (`get_token_balance` tells you how far your plan goes).
3. **Start.** When you say yes it starts the book. Each chapter is billed when it
starts rendering; a chapter that fails is refunded automatically and made
again on its own (see [failures](#stopping-resuming-failures)).
4. **The wait.** Audiobooks are made while VieNeu's servers are not busy with
people waiting for audio — expect hours rather than minutes, often overnight.
Books waiting at the same time take turns, a chapter each, so a short book
is never stuck behind a long one. You do not need to keep the chat open: when the book is done, a notification
appears under the bell on vieneu.io.
5. **Listening.** Ask *"Sách nói xong chưa?"* The assistant shows the progress
and, for each finished chapter, a download link (valid 24 hours — ask again
for fresh ones). In Claude and ChatGPT a player lists the chapters and plays
them one after another. Every chapter is also saved in your **Library** on
vieneu.io.
## Voices
- **The narrator** reads the narration and the chapter titles.
- **Characters** — up to 30 — each get their own voice. Minor characters can
stay with the narrator: the assistant writes their lines as narration.
- Describe what you want ("một bà cụ miền Nam", "cậu bé tám tuổi") and the
assistant searches the catalogue with `list_voices`. Your own cloned voices
cannot read audiobooks yet.
## How a chapter sounds
| What | How |
|---|---|
| Between paragraphs and lines | 0.6 s of silence |
| After the chapter title | 2.5 s |
| At a scene change | 2 s |
| Start and end of the file | 0.5 s of silence before, 3.5 s after |
| Loudness | −19 LUFS integrated, peaks at most −3 dBTP |
| File | MP3, 128 kbps, mono, 48 kHz, tagged with title, book, author and track number |
Stores differ in the files they accept — Audible, for one, asks for 192 kbps at
44.1 kHz — so check your store's requirements before uploading.
## Tools
| Tool | What it does | Cost |
|------|--------------|------|
| `create_audiobook` | Create the book: title, author, narrator, characters' voices | Free |
| `add_audiobook_chapter` | Add a chapter as a script in reading order, or continue a long one | Free¹ |
| `start_audiobook` | Start the book — also resumes a stopped book and redoes failed chapters | Tokens |
| `get_audiobook` | Progress, and the download link of each finished chapter | Free |
| `list_audiobooks` | Your books, newest first — finds a book from an earlier chat | Free |
| `cancel_audiobook` | Stop: chapters not started are never billed; one already being made finishes | Free |
¹ A chapter added to a book that is already being made joins it, and is billed
when it is made.
A chapter's script is a list of segments: narration (no speaker), or one
character's spoken words. A line such as *"— Đi thôi! — Lan nói."* becomes Lan
saying *"Đi thôi!"* followed by the narrator reading *"Lan nói."*
## Limits
| Limit | Value |
|---|---|
| Chapters per book | 200 |
| Characters per chapter | 60,000 (about 80 minutes), less on a plan with a daily cap¹ |
| Characters per book | 1,500,000 |
| Characters with their own voice | 30 |
¹ A chapter is billed in one go, so it has to fit the plan's daily (or weekly)
token cap — on a trial of 50,000 tokens a day that is about 12,800 characters.
The book reports its ceiling (`chapter_chars_max`), and a chapter over it is
refused when added, with the size that fits, rather than waiting for tokens that
never come. The assistant splits such a chapter into several.
An assistant writes every word of the book into its tool calls, a few thousand
characters at a time, so long chapters arrive in several pieces — that is
normal. Each piece says how many segments the chapter already has
(`after_segment`): a piece sent twice is refused instead of being added twice,
a piece that went missing is noticed before the next one lands, and each answer
shows the words the chapter now ends with. Before starting, the assistant checks
every chapter against your text. A whole novel is a long conversation: it is
fine to carry on in a new chat, where *"tiếp tục sách nói …"* finds the book again.
## Stopping, resuming, failures
- **"Dừng sách nói"** stops the book. Chapters not started are not billed; the
one being made finishes, because it is already paid for.
- **"Làm tiếp"** resumes it, with the cost of what is left.
- A chapter whose rendering fails — a server restarting mid-chapter, a voice
server down — is refunded and made again on its own a few minutes later,
continuing from the parts already made, twice. Only a third failure counts:
the book then ends as *partly done*, and starting it again redoes only the
failed chapters.
- A chapter whose job stops moving for 45 minutes is treated the same way, so
one stuck chapter never holds up the books behind it.
- If the finishing step (the MP3) fails, it is tried again with growing waits for
about an hour; only then is the chapter handed over as WAV instead.
## Your text
You are reading your text to VieNeu, so make sure you may: your own writing, a
public-domain work, or one you hold the audio rights to.
## From your own code
The same books are in the Cloud API: `POST /v1/audiobooks`, then
`POST /v1/audiobooks/{bookId}/chapters` and `POST /v1/audiobooks/{bookId}/start`;
follow progress with `GET /v1/audiobooks/{bookId}`. See
[Every operation](../../cloud-api/overview.md#every-operation) and the
[API reference](/api-reference).
---
# Use VieNeu in Claude
Source: https://docs.vieneu.io/docs/integrations/mcp/claude-ai
Works in Claude on the **web** (claude.ai), **Claude Desktop** and the **Claude
mobile** apps — one connector, added once, available everywhere you sign in to
Claude. The audio plays in a small [player](./index.md#the-inline-player)
right under the reply.
**You need:** a Claude account (Free plans can add **one** custom connector;
Pro and Max can add more) and a VieNeu account with an active token plan.
## Add the connector (Free, Pro, Max)
1. In Claude, open **Customize → Connectors**
([claude.ai/customize/connectors](https://claude.ai/customize/connectors)).
2. Click **Add custom connector** (on some screens: **+** first).
3. Fill in:
- **Name:** `VieNeu`
- **URL:** `https://api.vieneu.io/mcp`
Leave any OAuth client fields empty — Claude identifies itself to VieNeu on
its own.
4. Click **Add**, then **Connect**.
5. A VieNeu page opens. Sign in if asked, check that the **account** shown is
the one you want, and press **Cho phép** (Allow). You land back in Claude.
Add it on the web or in Claude Desktop; the mobile apps then use the same
connector.
## Add the connector (Team, Enterprise)
Only an organization **Owner** can add a custom connector.
1. The Owner opens **Organization settings → Connectors**, chooses **Add →
Custom** (then **Web** if asked), enters `https://api.vieneu.io/mcp` and
clicks **Add**.
2. Each member then opens **Customize → Connectors**, finds **VieNeu** (marked
*Custom*), clicks **Connect**, and approves their **own** VieNeu account on
the VieNeu page. Each member's usage is billed to their own VieNeu plan.
## Use it in a chat
1. In the chat box, click **+** → **Connectors** and make sure **VieNeu** is
switched on for this conversation.
2. Ask in plain language, for example:
- "Tìm giọng nữ miền Bắc rồi đọc câu: Xin chào, đây là VieNeu."
- "Đọc đoạn này bằng giọng Thu Trang, tốc độ 1.1."
3. The first time, Claude asks permission to use a VieNeu tool — choose
**Allow once** or **Always allow**. It also asks once before showing the
player — choose **Allow** (or **Always allow**).
Each `text_to_speech` call spends tokens. Claude may offer a couple of voices to
compare; each one it reads is a separate charge.
## Using an API key instead
Claude's connector dialog can send a fixed request header only for a limited set
of organizations (a beta feature, Owners only). If your organization has a
**Request headers** section when adding the connector, choose **No sign-in** and
add the header `x-api-key` with your key, or `authorization` with the value
`Bearer vn_sk_...`. The key is shared by everyone in the organization — use
sign-in unless you specifically need this.
## Remove it
- From Claude: **Customize → Connectors → VieNeu → Remove** (or disconnect).
- From VieNeu: **Developer → Dùng VieNeu trong Claude, ChatGPT → Gỡ kết nối**.
This signs Claude out immediately, even on other devices.
## Problems
| Symptom | Fix |
|---|---|
| "Couldn't reach the MCP server" when adding | Check the URL is exactly `https://api.vieneu.io/mcp` (no trailing slash, no spaces) |
| The VieNeu page says the request expired | Go back to Claude and click **Connect** again — a request is valid for 30 minutes |
| "Tài khoản chưa có gói token còn hạn" on the VieNeu page | Buy or renew a plan on vieneu.io, then connect again |
| No player, only a link | Claude asks once before showing it — choose **Allow**. The link always works |
| Claude says it has no VieNeu tools | Turn the connector on for the conversation (**+ → Connectors**) |
More in [Troubleshooting](./troubleshooting.md).
---
# Use VieNeu in ChatGPT
Source: https://docs.vieneu.io/docs/integrations/mcp/chatgpt
ChatGPT connects to custom MCP servers through **developer mode**. The audio
plays in the [inline player](./index.md#the-inline-player) under the reply.
**You need:** ChatGPT **Plus, Pro, Business, Enterprise or Edu**, on
**chatgpt.com** (the web). The Free plan cannot add custom apps. On Business
and Enterprise, a workspace admin may have to allow developer mode or custom
apps first. And a VieNeu account with an active token plan.
:::note
ChatGPT's menus for custom apps have been renamed several times. The steps below
match OpenAI's guide as of October 2026; if a label differs, look for
"Developer mode" and "add app / connector" in the same area.
:::
## Add the app
1. Open **Settings → Security and login** and turn on **Developer mode**.
2. Open **Plugins** (apps) and click **+** to create one.
3. Fill in:
- **Name:** `VieNeu`
- **Description:** `Vietnamese text-to-speech`
- **Server URL / Connection:** `https://api.vieneu.io/mcp`
- **Authentication:** **OAuth**
4. Create it. ChatGPT registers itself with VieNeu automatically and opens the
VieNeu page: sign in if asked, check the account and press **Cho phép**
(Allow).
5. ChatGPT lists the five VieNeu tools. The app appears under **Drafts**.
ChatGPT does not support API keys for custom apps — use sign-in.
## Use it in a chat
1. In a new chat, open **+ → Developer mode** and select **VieNeu**.
2. Ask, for example: *"Tìm giọng nam miền Nam đọc tin tức, rồi đọc đoạn sau…"*
3. ChatGPT asks you to confirm before a tool that makes something (like
`text_to_speech`) runs — confirm it. Each synthesis spends tokens.
## Remove it
- In ChatGPT: remove the app from **Plugins** (or turn developer mode off).
- On VieNeu: **Developer → Dùng VieNeu trong Claude, ChatGPT → Gỡ kết nối**.
## Problems
| Symptom | Fix |
|---|---|
| No "Developer mode" option | Your plan or workspace does not allow it — see *You need* above |
| The VieNeu page warns the app "calls itself ChatGPT" but goes elsewhere | Do **not** allow it: the request did not come from ChatGPT. Real ChatGPT returns to `chatgpt.com` |
| Tools are missing after an update | Open the app under **Drafts** and click **Refresh** |
More in [Troubleshooting](./troubleshooting.md).
---
# Use VieNeu in Claude Code
Source: https://docs.vieneu.io/docs/integrations/mcp/claude-code
## Add the server
```bash
claude mcp add --transport http vieneu https://api.vieneu.io/mcp
```
Then, inside Claude Code, run `/mcp`, pick **vieneu** and choose
**Authenticate**. Your browser opens the VieNeu consent page. It shows a yellow
note that the app runs on your computer (the callback is `127.0.0.1` or
`localhost`) — that is expected for Claude Code. Press **Cho phép** (Allow).
From a shell you can also start the sign-in with:
```bash
claude mcp login vieneu
```
### Where it is saved (`--scope`)
| Scope | Stored in | Use when |
|---|---|---|
| `local` (default) | `~/.claude.json`, this project only | Just you, just here |
| `user` | `~/.claude.json`, every project | You want VieNeu everywhere |
| `project` | `.mcp.json` in the repository | The whole team should get it (each person still signs in) |
```bash
claude mcp add --transport http --scope user vieneu https://api.vieneu.io/mcp
```
## Using an API key instead
For CI or machines where a browser sign-in is not possible:
```bash
claude mcp add --transport http vieneu https://api.vieneu.io/mcp \
--header "X-API-Key: vn_sk_..."
```
In a shared `.mcp.json`, reference an environment variable instead of the key:
```json
{
"mcpServers": {
"vieneu": {
"type": "http",
"url": "https://api.vieneu.io/mcp",
"headers": { "X-API-Key": "${VIENEU_API_KEY}" }
}
}
}
```
## Use it
Ask in the session, for example: *"Đọc README tiếng Việt này bằng giọng nữ miền
Nam và cho mình link tải."* Claude Code shows the reply and the download link;
it has no inline player.
## Remove it
```bash
claude mcp remove vieneu
```
and, if you signed in, **Developer → Gỡ kết nối** on vieneu.io.
---
# Use VieNeu in Cursor
Source: https://docs.vieneu.io/docs/integrations/mcp/cursor
## Add the server
**One click:** **Add VieNeu to Cursor**. Cursor opens
with the server filled in; confirm **Install**, then sign in (below).
Or by hand: create or edit `mcp.json` — **`.cursor/mcp.json`** in a project, or
**`~/.cursor/mcp.json`** for every project:
```json
{
"mcpServers": {
"vieneu": { "url": "https://api.vieneu.io/mcp" }
}
}
```
Save it. Cursor lists **vieneu** under its MCP settings and asks you to sign in:
your browser opens the VieNeu page — press **Cho phép** (Allow). The page notes
the app runs on your computer; that is expected for Cursor.
## Using an API key instead
```json
{
"mcpServers": {
"vieneu": {
"url": "https://api.vieneu.io/mcp",
"headers": { "X-API-Key": "${env:VIENEU_API_KEY}" }
}
}
}
```
`${env:VIENEU_API_KEY}` reads the key from your environment, so the file can be
committed without the secret.
## Use it
In Agent chat: *"Đọc đoạn mô tả sản phẩm này bằng giọng nữ miền Nam, trả link
MP3."* Cursor asks before running a tool (you can allow it). Results render in
the [inline player](./index.md#the-inline-player) where Cursor supports MCP
Apps; the link is always in the reply.
## Teams
On Cursor Business/Enterprise, admins can restrict which MCP servers members may
use (**Team Settings → MCP Configuration**). Ask your admin to allow
`https://api.vieneu.io/mcp`.
---
# Use VieNeu in VS Code with GitHub Copilot
Source: https://docs.vieneu.io/docs/integrations/mcp/vscode
## Add the server
**One click:** **Add VieNeu to VS Code**. VS Code opens
the server's install page; choose **Install**, then sign in when it starts.
Or run **MCP: Add Server** from the Command Palette and choose **HTTP**, or create
**`.vscode/mcp.json`** in your workspace (or run **MCP: Open User
Configuration** for every workspace):
```json
{
"servers": {
"vieneu": {
"type": "http",
"url": "https://api.vieneu.io/mcp"
}
}
}
```
The first time the server starts, VS Code asks whether you trust it, then opens
your browser for sign-in: press **Cho phép** (Allow) on the VieNeu page.
## Using an API key instead
VS Code can prompt for the key once and store it securely:
```json
{
"inputs": [
{ "type": "promptString", "id": "vieneu-key", "description": "VieNeu API key", "password": true }
],
"servers": {
"vieneu": {
"type": "http",
"url": "https://api.vieneu.io/mcp",
"headers": { "X-API-Key": "${input:vieneu-key}" }
}
}
}
```
## Use it
Open Copilot Chat in **Agent** mode, make sure the VieNeu tools are enabled in
the tools picker, and ask: *"Đọc changelog này bằng giọng nam miền Bắc."*
**Inline player:** VS Code shows MCP Apps behind an experimental setting. Turn
on `chat.mcp.apps.enabled` to get the [player](./index.md#the-inline-player) in
chat; otherwise you get the link.
## Copilot Business and Enterprise
Organizations must enable the **"MCP servers in Copilot"** policy (off by
default) before members can use any MCP server. Copilot Free, Pro and Pro+ are
not affected.
---
# Use VieNeu from the OpenAI API
Source: https://docs.vieneu.io/docs/integrations/mcp/openai-api
If you build your own app on OpenAI models, you can hand the model the VieNeu
tools directly: the model decides when to search voices and synthesize, and
OpenAI calls `https://api.vieneu.io/mcp` for you.
**You need:** an OpenAI API key and a VieNeu [API key](../../cloud-api/overview.md#authentication)
(`vn_sk_...`, from the **Developer** page). Keep both in environment variables:
```bash
export OPENAI_API_KEY=sk-...
export VIENEU_API_KEY=vn_sk_...
```
:::tip Just want audio from code?
If your code already knows the text and the voice, you do not need MCP or an
LLM at all — call the [OpenAI-compatible endpoint](../../cloud-api/openai-compatible.md)
directly with the OpenAI SDK. MCP is for letting the *model* decide.
:::
## Responses API
The Responses API has a built-in remote MCP tool. Pass your VieNeu key in
`authorization`; OpenAI sends it to VieNeu as a Bearer token, which VieNeu
accepts.
```python
import os
from openai import OpenAI
client = OpenAI()
resp = client.responses.create(
model="gpt-5", # any current model that supports tools
tools=[{
"type": "mcp",
"server_label": "vieneu",
"server_url": "https://api.vieneu.io/mcp",
"authorization": os.environ["VIENEU_API_KEY"],
"allowed_tools": ["list_voices", "text_to_speech", "get_speech_status", "get_token_balance"],
"require_approval": "never",
}],
input="Tìm một giọng nữ miền Bắc rồi đọc câu: Xin chào, đây là VieNeu. Trả về link tải.",
)
print(resp.output_text)
```
```js
import OpenAI from "openai";
const client = new OpenAI();
const resp = await client.responses.create({
model: "gpt-5",
tools: [{
type: "mcp",
server_label: "vieneu",
server_url: "https://api.vieneu.io/mcp",
authorization: process.env.VIENEU_API_KEY,
allowed_tools: ["list_voices", "text_to_speech", "get_speech_status", "get_token_balance"],
require_approval: "never",
}],
input: "Tìm một giọng nữ miền Bắc rồi đọc câu: Xin chào, đây là VieNeu. Trả về link tải.",
});
console.log(resp.output_text);
```
- **`require_approval`:** the default asks for approval before every tool call
(the response then contains `mcp_approval_request` items you must answer).
`"never"` lets the model spend your VieNeu tokens without asking — keep
`allowed_tools` tight, or require approval for `text_to_speech` only:
`"require_approval": {"always": {"tool_names": ["text_to_speech"]}}`.
- **The key is not stored by OpenAI**, so send it with every request.
- The audio link is in the tool result and usually in the model's answer; the
`mcp_call` items in `resp.output` hold the raw tool results.
## OpenAI Agents SDK (Python)
Let OpenAI call VieNeu for you (hosted tool):
```python
import os
from agents import Agent, HostedMCPTool, Runner
agent = Agent(
name="Narrator",
instructions="You narrate Vietnamese text with VieNeu and return the audio link.",
tools=[HostedMCPTool(tool_config={
"type": "mcp",
"server_label": "vieneu",
"server_url": "https://api.vieneu.io/mcp",
"authorization": os.environ["VIENEU_API_KEY"],
"require_approval": "never",
})],
)
result = Runner.run_sync(agent, "Đọc câu 'Chào buổi sáng' bằng giọng Thu Trang.")
print(result.final_output)
```
Or connect from your own process (your code calls VieNeu, with any header you
choose):
```python
import asyncio, os
from agents import Agent, Runner
from agents.mcp import MCPServerStreamableHttp
async def main():
async with MCPServerStreamableHttp(
name="vieneu",
params={
"url": "https://api.vieneu.io/mcp",
"headers": {"X-API-Key": os.environ["VIENEU_API_KEY"]},
"timeout": 60,
},
cache_tools_list=True,
) as vieneu:
agent = Agent(name="Narrator", mcp_servers=[vieneu])
result = await Runner.run(agent, "Đọc câu 'Chào buổi sáng' bằng giọng Thu Trang.")
print(result.final_output)
asyncio.run(main())
```
Set `timeout` generously: `text_to_speech` waits for the audio (a few seconds
for short text, up to about 50 seconds for long text).
## Costs
Two bills: OpenAI charges for the model's tokens, VieNeu charges your plan for
each synthesis (per character, minimum 50). `list_voices`, `list_emotion_tags`,
`get_speech_status` and `get_token_balance` are free on the VieNeu side.
---
# Use VieNeu in other MCP apps
Source: https://docs.vieneu.io/docs/integrations/mcp/other-clients
Any app that supports **remote MCP servers over Streamable HTTP** works with
`https://api.vieneu.io/mcp`. Apps that support MCP sign-in open the VieNeu page
on their own; the rest can send an [API key](./index.md#using-an-api-key-instead-of-signing-in).
Replace `vn_sk_...` with your key, or better, keep it in an environment variable
`VIENEU_API_KEY` as shown.
## Windsurf (Devin Desktop)
Windsurf is now called **Devin Desktop**. Edit its `mcp_config.json` (open it
from the MCP settings in the app — the file's location changed with the
rename). Note the key is `serverUrl`, not `url`:
```json
{
"mcpServers": {
"vieneu": {
"serverUrl": "https://api.vieneu.io/mcp",
"headers": { "X-API-Key": "${env:VIENEU_API_KEY}" }
}
}
}
```
Leave out `headers` to sign in instead.
## Gemini CLI
```bash
gemini mcp add --transport http vieneu https://api.vieneu.io/mcp
```
Then sign in inside the CLI with `/mcp auth vieneu`. Or edit
`~/.gemini/settings.json` (`httpUrl` means Streamable HTTP):
```json
{
"mcpServers": {
"vieneu": {
"httpUrl": "https://api.vieneu.io/mcp",
"headers": { "X-API-Key": "vn_sk_..." }
}
}
}
```
## OpenAI Codex CLI
```bash
codex mcp add vieneu --url https://api.vieneu.io/mcp
codex mcp login vieneu
```
Or with an API key, in `~/.codex/config.toml`:
```toml
[mcp_servers.vieneu]
url = "https://api.vieneu.io/mcp"
bearer_token_env_var = "VIENEU_API_KEY" # sends Authorization: Bearer
```
## Zed
In Zed's `settings.json`:
```json
{
"context_servers": {
"vieneu": {
"url": "https://api.vieneu.io/mcp",
"headers": { "Authorization": "Bearer vn_sk_..." }
}
}
}
```
Without an `Authorization` header Zed starts the sign-in flow instead.
## Goose
`goose configure` → **Add Extension** → **Remote Extension (Streamable HTTP)**
→ URL `https://api.vieneu.io/mcp` (in Goose Desktop: **Extensions → Add custom
extension**). Goose signs in on its own, or you can add an `X-API-Key` header in
the wizard. Goose shows the [inline player](./index.md#the-inline-player).
## LM Studio
LM Studio uses a token header (no sign-in). In its `mcp.json`:
```json
{
"mcpServers": {
"vieneu": {
"url": "https://api.vieneu.io/mcp",
"headers": { "Authorization": "Bearer vn_sk_..." }
}
}
}
```
## An app not listed here
Look for "remote MCP server", "HTTP" or "Streamable HTTP" in its settings and
use the URL. If it only supports local (stdio) servers, bridge it with
[`mcp-remote`](https://www.npmjs.com/package/mcp-remote):
```json
{
"mcpServers": {
"vieneu": {
"command": "npx",
"args": ["-y", "mcp-remote", "https://api.vieneu.io/mcp"]
}
}
}
```
`mcp-remote` opens the VieNeu sign-in page in your browser the first time.
---
# MCP troubleshooting
Source: https://docs.vieneu.io/docs/integrations/mcp/troubleshooting
## Connecting
| Symptom | Cause | Fix |
|---|---|---|
| "Couldn't reach the MCP server" / connection failed | Wrong URL | Use exactly `https://api.vieneu.io/mcp` — no trailing slash, no `/api` |
| The VieNeu page says the request expired or was already handled | The sign-in page was left open more than 30 minutes, or approved twice | Go back to the app and connect again |
| "Tài khoản chưa có gói token còn hạn" on the VieNeu page | No active plan | Buy or renew a plan on vieneu.io, then connect again |
| "Tài khoản đã có 20 ứng dụng kết nối" | Too many connections | Remove old ones under **Developer → Gỡ kết nối** |
| The VieNeu page signs in a different account than you expected | You were already signed in to vieneu.io in that browser | Click **Đổi tài khoản** on the page |
| A red warning: the app "calls itself Claude/ChatGPT" but returns elsewhere | The request did not come from that app | Press **Từ chối** (Deny). Real Claude returns to `claude.ai`, ChatGPT to `chatgpt.com` |
| A yellow note: the app runs on your computer | Normal for Claude Code, Cursor, VS Code and other desktop tools | Allow it if you just clicked connect in that tool |
## Using it
| Symptom | Cause | Fix |
|---|---|---|
| The assistant says it has no VieNeu tools | The connector is off for this chat | Turn it on (Claude: **+ → Connectors**; ChatGPT: **+ → Developer mode**) |
| New VieNeu features (audiobooks, "what can VieNeu do?") don't show up | The app keeps the list of tools it fetched when you connected | Disconnect VieNeu and connect again, then start a new chat (Claude: **Settings → Connectors → VieNeu → Disconnect**, then **Connect**). Apps that load their tools when they start (Claude Code, Cursor, VS Code) only need a restart |
| Suddenly asked to sign in again | Removed on vieneu.io, or the sign-in expired | Reconnect from the app |
| "Voice … is not available" | The assistant guessed a voice id | Ask it to search with `list_voices` first |
| "Tài khoản không đủ quyền hoặc hết token…" | Out of tokens or plan expired | Top up or renew — retrying does not help |
| "Audio đang được tạo…" with a job id | Long text still rendering | Ask the assistant to check again in a moment (it uses `get_speech_status`, free) |
| No player, only a link | The app does not show MCP Apps (Claude Code and other CLIs), or you declined it | Use the link; in Claude choose **Allow** when asked to show the app |
| The link stopped working | Links expire after 24 hours | Download it sooner, or find the audio in your **Library** on vieneu.io |
## Charges
Every `text_to_speech` call is one synthesis, billed per character (minimum
50). Assistants sometimes offer to read the same sentence with a few different
voices — each of those is charged. A failed synthesis is refunded
automatically. See usage under **Developer → Usage** on vieneu.io.
## Still stuck?
Contact VieNeu support from vieneu.io with the time of the attempt and the app
you used. For developers: the server's discovery documents are at
[`/.well-known/oauth-protected-resource`](https://api.vieneu.io/.well-known/oauth-protected-resource)
and [`/.well-known/oauth-authorization-server`](https://api.vieneu.io/.well-known/oauth-authorization-server).
---
# Cloud API overview
Source: https://docs.vieneu.io/docs/cloud-api/overview
The VieNeu Cloud API turns Vietnamese text into speech over HTTPS. You send text
and a voice id, you get audio back — no model to download, no GPU to rent.
:::note Two different products
This section documents the **hosted API** at `api.vieneu.io`. The **SDK** section
documents the separate on-device package, which runs a model on your own machine.
They share voices and a name, and nothing else: different install, different
billing, different code. If you are integrating VieNeu into a product, you almost
certainly want this section.
:::
If you would rather have audio in front of you before reading any of this, the
**[Quickstart](./quickstart.md)** is a key, a voice and one call.
## The full reference
Every endpoint, every field, every error is in the **[API reference](/api-reference)**,
rendered from the OpenAPI specification the server itself generates. No login
required.
The same document is served as a plain file at
**[`/openapi.json`](pathname:///openapi.json)** — point `openapi-generator`,
`oapi-codegen`, Kiota or your language's equivalent at it and you have a typed
client:
```bash
curl -O https://docs.vieneu.io/openapi.json
```
It is generated from the API's own route definitions, and CI fails the build if
the committed file no longer matches them — so the spec cannot quietly fall
behind the *declarations* in the code.
That is a narrower guarantee than it sounds, and worth stating plainly: the check
compares the committed JSON against the decorators, not against what the server
does. A decorator that describes an endpoint incorrectly ships a spec that is
wrong and a check that is green — which is exactly how the cloning endpoint went
on advertising an engine choice for months after the server started refusing one.
Where this documentation and the spec disagree about behaviour, the pages here
are the ones written against the running code.
It covers `/api/v1` only — the application's own dashboard, billing and
administration endpoints are not part of the public contract.
## Base URL
```
https://api.vieneu.io/api/v1
```
## Authentication
Every request carries your API key, either way round:
```bash
-H "Authorization: Bearer vn_sk_..." # or
-H "X-API-Key: vn_sk_..."
```
Keys beginning `vn_sk_` are live; `vn_test_` keys work the same way but cap each
request at 100 words. Create either on the
**[Developer page](https://www.vieneu.io/#/developer?create=test)** in the VieNeu web app — the
plaintext is shown once, at creation, and never again.
A key is always issued against an **active token grant**, so a new account has to
start one before the page will mint anything; without a grant, creation answers
`403` with `No active API token grant found.` The Developer page offers a free
7-day trial that starts one, once per account.
## The two ways to synthesize
**Synchronous** — you get the audio bytes in the response:
```bash
curl https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "input": "Xin chào, đây là VieNeu.", "voice": "Ngọc Lan" }' \
--output speech.mp3
```
This is the [OpenAI-compatible endpoint](./openai-compatible.md) — the fastest way
in if you already have an OpenAI client, and the one to use for short text.
**Asynchronous** — for long text, submit a job and poll it:
```bash
curl -X POST https://api.vieneu.io/api/v1/tts \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "text": "…", "voiceId": "Ngọc Lan" }'
# → { "jobId": "…", "status": "queued" }
# or, when that exact audio already exists (a cache hit):
# { "jobId": "…", "status": "completed", "audioUrl": "https://…" } — nothing to poll
curl https://api.vieneu.io/api/v1/tts/ \
-H "Authorization: Bearer $VIENEU_API_KEY"
# → { "status": "completed", "audioUrl": "https://…" }
```
With several jobs in flight, poll them together — one request, up to 50 ids,
each entry in the same shape as the single poll:
```bash
curl "https://api.vieneu.io/api/v1/tts?ids=,," \
-H "Authorization: Bearer $VIENEU_API_KEY"
# → { "jobs": [ { "jobId": "…", "status": "completed", "audioUrl": "…" }, … ],
# "missing": [] }
```
Ids that are unknown or not yours land in `missing`; the batch is never refused
because of one bad id. Prefer the batch form whenever you have more than one job
queued: the edge proxy caps concurrent connections per source address, and a
batch costs the server one database read where each single poll costs two.
And when you need audio to start playing before the whole text is generated, use
[streaming](./streaming.md).
### Retrying safely
`POST /v1/tts` creates the job and charges the tokens **before** it answers. If
your request times out, or the connection is reset while the answer is on its
way — a `502` from the edge during a deploy looks exactly like this — you cannot
tell "never happened" from "happened, answer lost", and a blind retry can pay
twice.
Send an `Idempotency-Key` header and it cannot. Choose the key yourself (8–128
printable ASCII characters, no whitespace; a UUID per job is the intended shape)
and reuse it on every retry of *that* request:
```bash
curl -X POST https://api.vieneu.io/api/v1/tts \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Idempotency-Key: 6f1c2e0a-7b3d-4c58-9e21-0a5d3b7f9c44" \
-H "Content-Type: application/json" \
-d '{ "text": "…", "voiceId": "Ngọc Lan" }'
```
- Same key, same body, first call finished → the **first** job comes back
(same `jobId`, nothing charged again) with `Idempotent-Replayed: true`.
- Same key, same body, first call still running → `409 IDEMPOTENCY_IN_FLIGHT`
with `Retry-After: 1`. Retry with the **same** key.
- Same key, different body → `422 IDEMPOTENCY_KEY_REUSED`. Keys are per request.
- A first call that failed with a 4xx (quota, unknown voice) releases its key,
so a corrected retry under the same key goes through.
Keys are scoped to your API key and remembered for 24 hours. Every keyed
response echoes the key back in `Idempotency-Key`. Without the header the route
behaves exactly as it always has.
## Every operation
Thirty-two operations on twenty-seven paths — the whole public surface. What each one
does, and nothing about its fields: the request bodies live in the
[API reference](/api-reference) alone, because a second description of the API is
a second thing to keep true, and the one that goes stale is never the generated
one.
The reference deep-links by operation id, so `operationId` below is also how you
jump straight to an entry — `/api-reference#operation/createTtsJob` — and how you
name it to a generated client.
**Synthesis**
| Operation | `operationId` | What it does |
|---|---|---|
| `POST /v1/audio/speech` | `createSpeech` | OpenAI-compatible synthesis, sync or streaming |
| `POST /v1/tts` | `createTtsJob` | Submit an asynchronous job |
| `GET /v1/tts/{jobId}` | `getTtsJob` | Poll a job for status and a download URL |
| `GET /v1/tts?ids=…` | `getTtsJobs` | Poll up to 50 jobs in one call |
| `POST /v1/tts/stream` | `streamSpeech` | Low-latency framed streaming |
| `POST /v1/dialogue` | `createDialogue` | Multi-speaker dialogue in one call |
| `POST /v1/dub` | `createDub` | Re-voice an existing recording |
| `POST /v1/srt` | `createSrtDub` | Dub a subtitle file into a timecode-aligned track |
| `POST /v1/vapi/speech` | `vapiSpeech` | [Vapi](./vapi.md) custom-voice webhook — raw PCM for voice agents |
**Voices and metadata**
| Operation | `operationId` | What it does |
|---|---|---|
| `GET /v1/voices` | `listVoices` | The catalogue, plus your own cloned voices when authenticated. No key required |
| `GET /v1/audio/voices` | `listOpenAiVoices` | Ids only, for OpenAI-compatible clients. **Does** need a key |
| `GET /v1/engines` | `listEngines` | Live engines, features and billing multipliers. No key required |
| `GET /v1/emotion-tags` | `listEmotionTags` | Reading styles and inline cue tags, per engine. No key required |
**Usage**
| Operation | `operationId` | What it does |
|---|---|---|
| `GET /v1/balance` | `getBalance` | Tokens this key can spend right now, the plan it bills and its daily / weekly caps. Not billed |
| `GET /v1/usage` | `getUsage` | What this key (or your whole account) did over a window: calls, tokens, seconds of audio, per day / route / voice, errors |
**Cloning — web only**
Cloned voices are created in the web Studio at
[vieneu.io/#/clone](https://vieneu.io/#/clone), not through this API. Enrolment
is a guided job: the Studio denoises the clip, transcribes it, and lets you hear
the result before the voice is saved. A bad reference degrades every later
generation with that voice, not just one request, which is why the API no longer
offers the shortcut. `POST /v1/voices`, `DELETE /v1/voices/{voiceId}`,
`POST /v1/clone`, `POST /v1/upload` and `POST /v1/prepare` answer `410` with
`code: "CLONE_WEB_ONLY"` and bill nothing.
Using a clone through the API is unchanged: it appears in `GET /v1/voices` when
you send your key (`"kind": "cloned"`), and its `clone_…` id is an ordinary
`voiceId` on `POST /v1/tts` and `POST /v1/tts/stream`, at the same price as a
catalogue voice. See [Cloned voices](./streaming.md#cloned-voices).
**Webhooks** — see [Webhooks](./webhooks.md) for the event contract and a
signature verifier you can paste in.
| Operation | `operationId` | What it does |
|---|---|---|
| `POST /v1/webhooks` | `createWebhookEndpoint` | Register an endpoint. Returns the signing secret once |
| `GET /v1/webhooks` | `listWebhookEndpoints` | List your endpoints |
| `GET /v1/webhooks/{endpointId}` | `getWebhookEndpoint` | Read one endpoint |
| `DELETE /v1/webhooks/{endpointId}` | `deleteWebhookEndpoint` | Delete an endpoint |
| `GET /v1/webhooks/{endpointId}/deliveries` | `listWebhookDeliveries` | Recent delivery attempts, with the failure reason |
| `POST /v1/webhooks/{endpointId}/rotate-secret` | `rotateWebhookSecret` | Rotate the signing secret, old one live for 24 hours |
**Audiobooks** — a narrator and up to 30 character voices reading chapters of
script. Books are made in the background, one chapter at a time, while the GPU
fleet is idle: expect hours, not seconds. Books waiting together take turns, a
chapter each. Each chapter is billed per character
when it starts, refunded if it fails, mastered to MP3 and saved to the Library.
Chapters are not reported as `tts.job.*` webhook events — follow the book. The
same feature from a chat: [Audiobooks](../integrations/mcp/audiobooks.md).
| Operation | `operationId` | What it does |
|---|---|---|
| `POST /v1/audiobooks` | `createAudiobook` | Create a draft: title, narrator, cast. LIVE keys only |
| `GET /v1/audiobooks` | `listAudiobooks` | Your books, newest first |
| `GET /v1/audiobooks/{bookId}` | `getAudiobook` | Progress, and a download link (24 h) per finished chapter |
| `POST /v1/audiobooks/{bookId}/chapters` | `addAudiobookChapter` | Add a chapter as a script, or continue one not started (`appendTo`) |
| `POST /v1/audiobooks/{bookId}/start` | `startAudiobook` | Queue it; `402` when the plan cannot pay. Also resumes and redoes failed chapters |
| `POST /v1/audiobooks/{bookId}/cancel` | `cancelAudiobook` | Stop. Chapters not started are never billed |
### Voice ids differ by engine
A voice belongs to one engine, and since **2026-09-24** the cloud API has
exactly one: `v4`, whose ids are Vietnamese display names — `Ngọc Lan`, not a
slug. The retired `v3` catalogue used opaque `vieneu-…` slugs, and the two
catalogues shared almost no ids, so a voice copied from an old `v3` listing does
not resolve anywhere any more. It is rejected with `400` rather than silently
mapped to a `v4` voice — and that is now the most common cause of `400`s on the
synthesis routes for callers who hard-coded a voice before the retirement,
shipped, and never had a reason to look at the id again.
**Resolve ids at runtime, scoped to the engine you are about to use:**
```bash
curl -s "https://api.vieneu.io/api/v1/voices?engine=v4"
```
No key needed. `GET /v1/voices` with no filter lists the same `v4` catalogue
(plus your own `v4` clones when you send a key); `?engine=v3` answers `400`
with the retirement message quoted under [Engines](#engines).
## Engines
Requests carry an optional `engine`. Since **2026-09-24 08:00 (GMT+7)** the
only value the cloud API accepts is `v4`, which is also the default, so the
field can simply be omitted. `v3` was retired that morning: any request naming
it — on `POST /v1/tts`, `/v1/tts/stream`, `/v1/audio/speech`, `/v1/dialogue`,
`/v1/dub`, `/v1/srt`, or as `GET /v1/voices?engine=v3` — is refused with `400`
before anything is billed, and the message says why:
```
Engine "v3" was retired on 2026-09-24. Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues share no ids.
```
If you had pinned `v3`, the fix has three parts: send `"engine": "v4"` (or
nothing), take a voice id from `GET /v1/voices?engine=v4` rather than reusing
the old one, and re-enrol any clone you made on `v3` in the
[Studio](https://vieneu.io/#/clone). Audio you already generated on `v3` is
untouched — library entries and download links keep working.
The engine is billed at its own multiplier, applied on top of the per-character
rate:
| Engine | Multiplier |
|---|---|
| `v4` | 3× |
That figure is set per deployment and can change, so the table above is a
snapshot, not a contract. **`GET /v1/engines` returns the live values** — no API
key needed. If you are budgeting a large workload, read them from there rather
than from this page:
```bash
curl -s https://api.vieneu.io/api/v1/engines
```
```json
{ "globalMultiplier": 1.3,
"engines": [
{ "key": "v4", "isDefault": true, "sampleRate": 48000, "billingMultiplier": 3,
"features": ["generate", "clone", "dialogue", "stream", "dub", "srt"] }
] }
```
### The second multiplier
`globalMultiplier` in that response is a platform-wide lever sitting **on top of**
the per-engine rate. The full charge for one request is:
```
tokens = characters × engine multiplier × globalMultiplier
```
It spends long stretches at `1`, which is exactly why it is easy to miss — and
why it is returned alongside the engine rate rather than documented somewhere
else. At the time of writing it is `1.3`. Budget from the per-engine number alone
and your estimate is short by precisely this factor on any deployment where it
has been moved.
(A third factor, the AI text-refinement surcharge, applies only to requests that
ask for that step; it is not part of the base rate.)
`features` is also how you check, in code, what an engine can do — `stream` is
on `v4`'s list, so every stream runs there and is billed at the `v4` rate, the
same rate as every other call now. A voice belongs to exactly one engine —
`GET /v1/voices?engine=v4` lists that engine's catalogue. Passing a voice from
the retired `v3` catalogue is rejected rather than silently substituted.
**Streaming is `v4` only**, and was before the retirement, for a reason worth
keeping in mind if you ever compare providers: `v3` never really streamed — it
delivered its chunks in a burst at the end — so `POST /v1/tts/stream` refused it
rather than making a time-to-first-audio promise it could not keep. Now that
`v3` is gone from every route the gate is moot: a stream request naming `v3`
gets the same retirement `400` as any other request, and so does
`POST /v1/audio/speech` with `"model": "vieneu-v3"`, which until 2026-09-24
accepted the request, answered 200, and delivered the burst at the end with
nothing in the response saying so. If you are pinning an engine for a streaming
client, pin `v4` — see
[Drop-in for OpenAI-compatible apps](../integrations/openai-clients.md).
A cloned voice streams too, provided it is enrolled on `v4` — every clone
created since 2026-08-28 is, and an older `v3`-enrolled clone can no longer be
rendered at all until it is re-enrolled. See
[Cloned voices](./streaming.md#cloned-voices).
## Billing
Synthesis is billed per **submitted character**, so the cost of a call is
predictable from the request itself, before you make it. Minimum 50 characters
per request. A request that fails to produce audio is refunded automatically.
`aiRefine` defaults to **off** on the API. With it off your text is synthesized
exactly as sent. Turn it on (`"aiRefine": true`) for the AI pass the web app uses
— formulas, acronyms and mixed-in English read correctly, and the content is
checked — billed with a surcharge and one extra round-trip of latency.
Deterministic text preparation runs either way, so Vietnamese is pronounced
correctly regardless.
### The one endpoint that is not billed per character
`POST /v1/dub` is priced by the **output audio duration** — 65 tokens per second,
with a 150-token floor — and charged **only on success**. Dubbing runs speech
recognition before it runs synthesis, so the work tracks how long the recording
is, not how many characters the speaker happened to fit into it. The engine
multiplier still applies on top.
That makes dub the one call whose cost you cannot compute from the request body
before sending it. An upload is gated for affordability against its own duration
first, so a request you cannot pay for is refused before any GPU time is spent on
it.
### Seeing what you were billed for
`GET /v1/balance` is the quick question — *can I afford the next call?*
`remaining` is what one call can spend right now: the key's plan's tokens, or
less when a daily or weekly cap is lower. `plan` has the totals, each cap's
remainder and reset time (`null` when there is no such cap) and the expiry;
`totalRemainingTokens` adds up every active plan on the account. A synthesis
needs at least `50 × engine multiplier` tokens.
`GET /v1/usage` answers "am I being counted correctly" without a support ticket.
It reads the same per-call ledger the operators see: for the calling key
(default) or every key on the account (`scope=account`), over a window of up to
92 days (`from`, `to`; default the last 30 days), it returns totals — calls by
outcome, tokens, characters, seconds of audio, average and p95 generation time —
and breakdowns per day, per route, per voice and per error code.
A few things to know when reading it. `POST /v1/tts` rows start as `queued` and
become `ok` or `error` when the job finishes; a request served from the content
cache is `cache_hit` (billed, no generation time). `tokenCost` is what was
deducted **before** the call ran; `refunded: true` on a row means it was given
back. A call refused by the rate limiter never reaches the ledger — count those
from your own 429s. `requestCount` on your key's own record is something else:
how many times the key *authenticated*, polling included.
## Rate limits and request ids
There are **two** limiters in front of you, and they count different things. An
integration that only knows about one of them will misread half its 429s.
**The application limiter counts your API key.** This is the one that matches
what you bought: it keys on the key itself, so spreading calls across machines
does not buy you more and sharing an office network does not cost you any. Limits
are per route, and every route on `/v1` sets its own:
| Routes | Per minute |
|---|---|
| `/v1/audio/speech`, `/v1/tts`, `/v1/tts/stream`, `/v1/vapi/speech` | 300 |
| `/v1/dialogue` | 20 |
| `/v1/dub`, `/v1/srt` | 10 |
| Webhook management | 10–30, per operation |
| `GET /v1/voices`, `/v1/audio/voices`, `/v1/engines`, `/v1/emotion-tags`, `GET /v1/tts/{jobId}`, `GET /v1/tts?ids=` | not throttled |
When this limiter refuses you, it says so:
```
X-RateLimit-Limit: 300
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 43
Retry-After: 43
X-Request-Id: 0f7c…
```
`X-RateLimit-Reset` is when the window rolls; `Retry-After` is when you may call
again. They usually agree, and when they do not, honour `Retry-After`.
**The edge proxy counts your source address.** In front of the application sits
nginx, which can only key on an IP. On `/api/v1` it allows **3000 requests per
minute and 300 concurrent connections per address**. That is deliberately loose —
it is DDoS padding, not a product limit — but it is real, it is **shared with
every other caller behind your address**, and it knows nothing about your key.
One customer calling from a dozen machines looks like a dozen customers there,
and a dozen customers behind one integration platform's egress IP look like one.
**The tell is the headers.** A `429` carrying `X-RateLimit-*` came from the
application and is about your key. A bare `429` or `503` with none of those
headers and no `Retry-After` came from the edge and is about your IP — nginx
answers with its own page and none of our headers. That difference decides what
to do: back off per key in the first case; in the second, back off *and* look at
how many machines share that egress address, because widening concurrency will
make it worse.
Streaming is the case where this matters most, because `/v1/tts/stream` matches a
different nginx location than the rest of `/api/v1` and inherits a much tighter
connection cap — see [Limits](./streaming.md#limits).
Every response, success or failure, carries `X-Request-Id`. Quote it in a support
request — it is the id our logs are keyed by.
## Errors
Native endpoints return the usual shape:
```json
{ "statusCode": 403, "message": "Insufficient tokens", "traceId": "0f7c…" }
```
`/v1/audio/speech` returns OpenAI's shape instead, so OpenAI clients can parse it.
| Status | Meaning |
|---|---|
| 400 | Malformed request — an unknown voice, a bad format, text the validator refused |
| 401 | Missing, malformed or revoked API key |
| 403 | Out of tokens, grant expired, or your plan does not include this engine |
| 422 | Content refused by moderation (only when `aiRefine` is on) |
| 429 | Rate limit or token quota — see above for telling them apart |
| 503 | No worker available for the requested engine or format; retry shortly |
Which refusals carry a machine-readable `code`, which do not, and what to do
about each is on the [Errors](./errors.md) page.
---
# Quickstart
Source: https://docs.vieneu.io/docs/cloud-api/quickstart
Three steps, and the third one returns audio. Nothing here needs an SDK.
## 1. Get an API key
Keys live on the **[Developer page](https://www.vieneu.io/#/developer?create=test)** in the
VieNeu web app. Create one, copy it once — the plaintext is shown at creation and
never again — and put it in your environment:
```bash
export VIENEU_API_KEY="vn_sk_..."
```
A key can only be issued against an **active token grant**, so a brand-new
account has to start one first; the Developer page offers a free 7-day trial that
does exactly that, once per account. Without a grant, key creation answers `403`
with `No active API token grant found.`
Keys beginning `vn_sk_` are live. `vn_test_` keys behave identically — same
billing, same limits, same endpoints — except that each request is capped at 100
words, which makes them safe to paste into a sample repository.
## 2. Pick a voice
The catalogue is public — this call needs no key:
```bash
curl -s "https://api.vieneu.io/api/v1/voices?engine=v4" | head -c 400
```
```json
{"voices":[{"id":"Ngọc Lan","description":"Giọng nữ, giọng trầm dịu dàng",
"name":"Ngọc Lan","gender":"female","region":"south","engine":"v4",
"kind":"catalog"}, …]}
```
The `id` field is what you send. **Always filter by engine.** A voice belongs to
one engine, the two catalogues are only partly interchangeable, and sending an id
from the wrong one is a `400` rather than a substitution — see
[voice ids differ by engine](./overview.md#voice-ids-differ-by-engine).
## 3. Synthesize
`POST /v1/audio/speech` returns the audio bytes in the response — no job, no
polling. It is OpenAI's endpoint shape, so an OpenAI client works against it
unchanged; see [OpenAI-compatible endpoint](./openai-compatible.md).
### curl
```bash
curl https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "input": "Xin chào, đây là VieNeu.", "voice": "Ngọc Lan" }' \
--output speech.mp3
```
### Python
```python
# pip install requests
import os
import requests
resp = requests.post(
"https://api.vieneu.io/api/v1/audio/speech",
headers={"Authorization": f"Bearer {os.environ['VIENEU_API_KEY']}"},
json={"input": "Xin chào, đây là VieNeu.", "voice": "Ngọc Lan"},
timeout=120,
)
resp.raise_for_status()
with open("speech.mp3", "wb") as f:
f.write(resp.content)
print("wrote speech.mp3", len(resp.content), "bytes")
```
### JavaScript
```js
// Node 18+ — no dependencies.
import { writeFile } from 'node:fs/promises';
const resp = await fetch('https://api.vieneu.io/api/v1/audio/speech', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.VIENEU_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({ input: 'Xin chào, đây là VieNeu.', voice: 'Ngọc Lan' }),
});
if (!resp.ok) throw new Error(`HTTP ${resp.status}: ${await resp.text()}`);
await writeFile('speech.mp3', Buffer.from(await resp.arrayBuffer()));
console.log('wrote speech.mp3');
```
You now have an mp3. `response_format` changes that — `wav`, `opus`, `pcm` and
8 kHz `ulaw` are all available, with the field-level detail in the
[API reference](/api-reference) under `createSpeech`.
## Where to go next
- **Long text** — `/v1/audio/speech` holds the connection open for the whole
synthesis. Past a few paragraphs, submit a job instead: see
[the two ways to synthesize](./overview.md#the-two-ways-to-synthesize).
- **Playback before the text finishes** — [Streaming](./streaming.md).
- **When a call fails** — [Errors](./errors.md) lists every machine-readable
`code` and what to do about it.
- **Everything else** — the [API reference](/api-reference) is generated from the
server's own route definitions and covers every field of every endpoint.
---
# OpenAI-compatible TTS endpoint
Source: https://docs.vieneu.io/docs/cloud-api/openai-compatible
VieNeu's native public API is asynchronous (`POST /v1/tts` returns a `jobId` you
poll). `POST /v1/audio/speech` is a **drop-in for OpenAI's endpoint of the same
name**: point an OpenAI SDK at VieNeu's base URL, give it your VieNeu key, and it
works without any other change.
This matters more than it looks. A large amount of software already speaks this
one endpoint — Open WebUI, SillyTavern, LobeChat, LiteLLM, LiveKit's and
Pipecat's OpenAI plugins, a long tail of scripts — and all of them accept a
custom `base_url`. Supporting this shape is what makes VieNeu usable in them with
no plugin, no adapter and no work on our side.
## Endpoint
```
POST /api/v1/audio/speech
Authorization: Bearer # vn_sk_... (X-API-Key also accepted)
Content-Type: application/json
```
Returns the audio **bytes synchronously** — no polling.
## Request body
| Field | OpenAI | VieNeu behavior |
|-------|--------|-----------------|
| `input` | required | the text to synthesize ✅ |
| `model` | required | selects the engine — and since 2026-09-24 there is one, `v4`. OpenAI's names (`tts-1`, `tts-1-hd`, `gpt-4o-mini-tts`), `vieneu` and `vieneu-v4` all resolve to it. Any other value is accepted and ignored, as before — except a `vieneu-…` name for an engine that does not exist or has been retired, which is rejected: `vieneu-v3` answers 400 with `Model 'vieneu-v3' is retired. Engine "v3" was retired on 2026-09-24. …`. |
| `voice` | `alloy`, … | a **VieNeu** voice id — list them with `GET /v1/audio/voices`. OpenAI's voice names are not mapped. Omit for the default voice. |
| `response_format` | `mp3` (default) | `mp3` (default), `wav`, `opus`, `pcm`, plus VieNeu's `ulaw`. `aac` and `flac` return 400. |
| `speed` | 0.25–4.0 | accepted across OpenAI's full range and **clamped** to 0.5–2.0, where the engine holds quality. A valid OpenAI value never returns an error. |
| `stream_format` | `audio` \| `sse` | supported for `pcm` and `ulaw` — see [Streaming](#streaming). |
| `instructions` | (gpt-4o-mini-tts) | accepted and ignored. Use inline cue tags instead. |
| `sample_rate` | — | VieNeu extension: 8000, 16000, 22050, 24000, 44100 or 48000. |
| `emotion` | — | VieNeu extension left over from `v3`: `natural` (default) or `storytelling`. Accepted and **ignored** on `v4` — see below. |
| `aiRefine` | — | VieNeu extension, **default `false`** — see [Billing](#billing). |
:::caution `emotion` does nothing any more
`emotion` was a `v3` parameter, and `v3` was retired from the cloud API on
2026-09-24. `v4` — now the only engine — has **no reading styles**. It renders
three inline cue tags — `[cười]`, `[thở dài]` and `[hắng giọng]` — and
**deletes every other tag from your text** rather than voicing it. A request
that sets `"emotion": "storytelling"` is accepted, billed, and read in the
ordinary voice; it is not an error, so nothing in the response tells you the
field went nowhere. Naming `"model": "vieneu-v3"` to get the old behaviour back
is a 400.
`GET /v1/emotion-tags` returns what the engine can actually render, and it needs
no API key. With no `engine` it answers for the default — `v4` — so the reply is
the three cues above and an empty list of styles; `?engine=v3` is a 400. Ask it
rather than assuming, especially if you are carrying a tag list written against
`v3`.
:::
### Formats
`pcm` and `ulaw` are **headerless**: the bytes carry no sample rate, so read it
from the `X-Sample-Rate` response header.
Rates: `pcm` defaults to **24000**, matching what OpenAI documents so a client
following their contract plays it at the right speed. Everything else defaults to
48000. `opus` is always 48 kHz (Opus itself is) and `ulaw` always 8 kHz, which is
what a phone line wants — passing a conflicting `sample_rate` is rejected rather
than quietly ignored.
> mp3 and opus need `ffmpeg` on the worker that serves the request. A node that
> predates it answers **503** naming the format, rather than returning something
> that is not the format you asked for. The one exception is a request that never
> named a format: since mp3 is *our* default rather than your choice, those fall
> back to wav instead of failing.
## Examples
**curl**
```bash
curl https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu.", "voice": "Ngọc Lan" }' \
--output speech.mp3
```
**OpenAI Python SDK** (unmodified, pointed at VieNeu)
```python
from openai import OpenAI
client = OpenAI(api_key="vn_sk_...", base_url="https://api.vieneu.io/api/v1")
resp = client.audio.speech.create(
model="tts-1",
voice="Ngọc Lan", # a VieNeu voice id, from GET /v1/audio/voices
input="Xin chào, đây là VieNeu.",
)
resp.stream_to_file("speech.mp3")
```
**Telephony (8 kHz mu-law)**
```bash
curl https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "input": "Xin chào", "response_format": "ulaw" }' \
--output line.raw # raw G.711 mu-law, 8000 Hz, no header
```
## Streaming
Set `stream_format` to start receiving audio while it is still being generated,
instead of waiting for the whole file.
- **`audio`** — the bytes arrive as chunked transfer encoding. Append them; the
result is the same audio you would have got in one piece.
- **`sse`** — Server-Sent Events. `speech.audio.delta` events carry base64 audio;
the stream ends with exactly one `speech.audio.done` (the audio is complete) or
`speech.audio.error` (it is not).
```python
with client.audio.speech.with_streaming_response.create(
model="tts-1", voice="Ngọc Lan", input="…", response_format="pcm",
extra_body={"stream_format": "audio"},
) as resp:
resp.stream_to_file("speech.pcm") # raw s16le, 24 kHz
```
**`speech.audio.done` is the only proof the audio is whole.** If a worker dies
mid-generation the stream stops, and a truncated stream is otherwise
indistinguishable from a short one. When that event does not arrive, discard the
audio — the request is refunded automatically.
### Streaming is `pcm` and `ulaw` only
Streamed audio is generated in chunks and each chunk is encoded independently, so
joining them only produces a valid result for the headerless formats:
| Format | Streamable | Why not |
|---|---|---|
| `pcm`, `ulaw` | ✅ | raw samples; concatenation *is* playback |
| `wav` | ❌ | every chunk repeats its 44-byte header mid-file |
| `mp3` | ❌ | each chunk re-applies the encoder's delay — measured at ~24 ms of inserted silence per seam, plus a click |
| `opus` | ❌ | each chunk is a complete Ogg stream; most browsers play only the first one |
Asking for a non-streamable format with `stream_format` returns 400 rather than
shipping audio with gaps in it. For a complete mp3 or opus **file**, drop
`stream_format` — the synchronous response has none of these problems.
This will widen once the worker can hold a single encoder open across a whole
stream; today it starts a fresh one per chunk.
## Voices
```
GET /api/v1/audio/voices[?engine=v4] → { "voices": ["Ngọc Lan", …] }
```
A companion to this endpoint for OpenAI-compatible clients, which look for this
route to fill their voice picker. `GET /v1/voices` is the richer version — names,
gender, region, and your own cloned voices.
## Errors
Returned in OpenAI's shape:
```json
{ "error": { "message": "...", "type": "invalid_request_error", "param": "input", "code": null } }
```
- `400 invalid_request_error` — missing `input`, an unknown or retired
`vieneu-…` model (`vieneu-v3`, since 2026-09-24), an unsupported
`response_format`, `stream_format` with a format that cannot be streamed, or
a `sample_rate` that contradicts the format.
- `401 authentication_error` — the key is missing, malformed or revoked.
- `403 insufficient_quota` — the key's token grant is exhausted or expired.
**Not 402:** nothing on `/v1` answers `402` at all. Being out of credit splits
across two statuses here, and the split is the part worth branching on: an
exhausted or expired grant is `403 insufficient_quota` and **will not clear on
its own**, while a spent daily or weekly cap is `429 rate_limit_exceeded` and
will. So `insufficient_quota` means top up or renew; `rate_limit_exceeded`
means wait.
- `422 content_policy_violation` — refused by moderation (only when `aiRefine` is on).
- `429 rate_limit_exceeded` — the rate limiter throttled you, **or** a daily or
weekly token cap is spent. This envelope cannot tell those two apart: the
native routes name the cap in `code` and say when it lifts in `resetAt`, and
neither field survives the translation into OpenAI's shape. See
[Errors](./errors.md#the-openai-compatible-route).
- `503` — no worker for the requested engine, or none that can encode the
requested format.
- `500 api_error` — synthesis failed (the token charge is **refunded**).
One caveat on `403`. The quota `403` above is typed by the handler itself. A
`403` raised *before* the handler runs — a guard refusing a feature your plan
does not include — is typed from the status alone, which folds `401` and `403`
together as `authentication_error`. So `insufficient_quota` always means money,
but `authentication_error` can mean either the key or the plan.
Every response carries `X-Request-Id`; quote it in a support request. The body
does not repeat it — OpenAI's envelope has no field for it, so read the header.
## Billing
Tokens are deducted from the API key's grant by the **submitted** character
count, so the cost of a call is predictable from the request alone. A failed
synthesis is refunded automatically.
`aiRefine` defaults to **`false`** here, unlike the web app, where it is on. With
it off the text is synthesized as submitted: no AI moderation, no pronunciation
normalization, no surcharge, and one less model round-trip of latency. Set it to
`true` to get the web app's behaviour — formulas, acronyms and mixed-in English
read correctly, content checked — billed with the AI surcharge.
Deterministic text preparation (chemistry spelling, ALL-CAPS folding, the sea-g2p
phoneme layer) runs either way. `aiRefine` controls only the AI call.
## Differences from OpenAI
- `voice` expects a VieNeu voice id, not an OpenAI voice name.
- `aac` and `flac` are not supported; `ulaw` and `sample_rate` are additions.
- `stream_format` covers `pcm` and `ulaw` only, not every format.
- `speed` outside 0.5–2.0 is clamped rather than honoured exactly.
- The richer VieNeu features — multi-speaker `dialogue`, `dub`, SRT dubbing —
have no OpenAI equivalent. Use the native `/v1/*` endpoints. Cloned voices are
made in the web Studio and then used as an ordinary `voiceId` on `POST /v1/tts`
(not on this route — a `clone_…` id here is a 400).
---
# Streaming
Source: https://docs.vieneu.io/docs/cloud-api/streaming
Long text takes as long to synthesize as it takes. Streaming lets playback start
on the first sentence — typically **1–2 seconds** on the standard plans, or a
ceiling agreed per contract on Enterprise — instead of after the last one.
:::info Engine — and what it costs
Streaming runs on **`v4`**. It always did — `v3` delivered its chunks in a
burst at the end, so a stream request naming it was refused rather than served
under a time-to-first-audio promise it could not keep — and since
**2026-09-24** `v4` is the only engine on the cloud API at all, so there is
nothing to choose: omit `engine`, or send `v4`. Naming `v3` on this route now
gets the same `400` as on every other: `Engine "v3" was retired on 2026-09-24.
Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues
share no ids.`
**The price is the `v4` price.** A stream is billed per submitted character at
`v4`'s multiplier — today **3×**, times the platform-wide `globalMultiplier` —
exactly like a `POST /v1/tts` job for the same text. The old comparison, a
stream at `v4`'s rate against a job on `v3`'s 1.5×, no longer means anything,
because the 1.5× rate is gone from every route; if you sized a budget against
it, resize it. `GET /v1/engines` returns the live multipliers, no key required
— see [Engines](./overview.md#engines).
This applies to cloned voices as well. Clones have enrolled on `v4` since
2026-08-28, so they stream like any other `v4` voice; an older `v3`-enrolled
clone has to be re-enrolled first. See [Cloned voices](#cloned-voices).
:::
There are two streaming endpoints, and which you want depends on who is reading
the bytes:
- **`POST /v1/audio/speech` with `stream_format`** — plain audio or SSE, and what
an OpenAI client already knows how to read. Start here.
- **`POST /v1/tts/stream`** — VieNeu's native framed stream. Use it when you want
each chunk delivered as a separately decodable unit, or when you want a
positive signal that the stream finished.
## The native stream
```bash
curl -N -X POST https://api.vieneu.io/api/v1/tts/stream \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "text": "…", "voiceId": "Ngọc Lan" }' \
--output stream.bin
```
The response body is a sequence of length-prefixed frames:
```
[4-byte big-endian uint32 = N][N bytes of audio] … repeated …
[4-byte big-endian uint32 = 0] ← end of stream
```
Read a length, read that many bytes, repeat. By default each payload is a
**self-contained WAV**, so you can hand a frame straight to a decoder without
waiting for the rest.
Response headers tell you what you actually got:
| Header | Meaning |
|---|---|
| `X-Sample-Rate` | Sample rate of the audio, in Hz |
| `X-Output-Format` | The encoding: `wav`, `mp3`, `opus`, `pcm` or `ulaw` |
| `X-Stream-Format` | `len32-wav-chunks`, or `len32-frames` for other encodings |
### The zero-length frame is the point
A stream that finishes sends a final frame with length `0`. A stream cut short by
a failure mid-generation simply **stops**, with no such frame.
Its absence means the stream was cut short: discard the audio. You are not
billed for a stream that never sent it — the refund is automatic.
**But the marker alone is not success.** A stream can arrive complete and still
carry no speech — a generation that produced nothing sends its heartbeats (see
below) and then a clean end-of-stream. We do not bill those either, so a caller
who treats the marker as sufficient will book a success, save a silent file, and
end the month reconciling against an invoice that never charged for it.
The condition we bill on, and the one you should use, is both halves:
> **the end-of-stream marker arrived, AND at least one frame carried samples.**
### Heartbeat frames
A stream that goes quiet for a while sends a **heartbeat**: a valid WAV file
containing **zero samples**, 44 bytes on the wire. Its only job is to keep the
connection from being closed for inactivity by whatever sits between us — a
proxy, a corporate gateway, a mobile carrier's NAT.
Heartbeats appear **only on WAV streams** (`X-Stream-Format: len32-wav-chunks`),
at most one every `X-Stream-Heartbeat` seconds, and only while the synthesizer
has produced nothing new. A stream that flows normally never sends one.
**You must skip them.** The trap is that a heartbeat is a *well-formed* WAV, so a
decoder will not reject it — `decodeAudioData` returns a zero-length buffer
rather than throwing, and writing the frame to a file leaves a stray 44-byte
header in the middle of your audio. Checking that a frame is non-empty is not
enough; ask whether it carries samples:
```python
def has_samples(frame: bytes) -> bool:
"""False for a heartbeat: a valid WAV whose `data` chunk is empty."""
if not frame:
return False # nothing there at all
if len(frame) < 12 or frame[:4] != b"RIFF" or frame[8:12] != b"WAVE":
return True # not a WAV we can read — assume audio
off = 12
while off + 8 <= len(frame):
chunk_id = frame[off:off + 4]
size = int.from_bytes(frame[off + 4:off + 8], "little")
if chunk_id == b"data":
return size > 0
off += 8 + size + (size % 2) # RIFF chunks are word-aligned
return True
```
Walk the chunks rather than assuming the `data` chunk starts at byte 36 — a
44-byte header is the common case, not a rule.
When you cannot read the header, treat the frame as audio. Guessing wrong in
that direction costs you one odd frame; guessing wrong in the other throws away
speech your listener was waiting for.
Raw codec streams (`pcm`, `ulaw`, `mp3`, `opus` — `X-Stream-Format:
len32-frames`) never carry heartbeats, because injecting a fake WAV into a codec
bitstream would be injecting garbage. If you stream those, every frame is audio.
### Formats
Pass `outputFormat` (and optionally `sampleRate`) to change the payload encoding:
```json
{ "text": "…", "voiceId": "Ngọc Lan", "outputFormat": "ulaw" }
```
| Format | Notes |
|---|---|
| `wav` | Default. Each frame is a self-contained file, 48 kHz. |
| `pcm` | Raw signed 16-bit little-endian, **no header** — read the rate from `X-Sample-Rate`. |
| `ulaw` | Raw G.711 mu-law, always 8 kHz. What a phone line wants. |
| `mp3`, `opus` | Available, but see the warning below. |
Valid `sampleRate` values are 8000, 16000, 22050, 24000, 44100 and 48000. The
framing never changes, whatever the encoding — so the zero-length terminator
means the same thing in all of them.
:::warning mp3 and opus frames do not join cleanly
Every frame is encoded as a standalone file, so **decode each frame separately**
— do not concatenate them.
Concatenated mp3 gains about 24 ms of silence at each seam (the encoder delay,
re-applied per frame) plus an audible click. Concatenated opus is a chain of
complete Ogg streams: ffmpeg reads it, most browsers stop at the first frame.
If you want one continuous mp3 or opus body rather than frames, request the whole
file without streaming — the synchronous path encodes it in one pass and has none
of these artifacts. `/v1/audio/speech`'s `stream_format` therefore accepts only
`pcm` and `ulaw`.
:::
## How a stream ends
Five outcomes, and they are not all failures. Your integration should tell them
apart, because three of them mean "try again" and two do not.
| What you see | What it means | Billed? |
|---|---|---|
| Frames, then a zero-length frame, at least one frame carrying samples | Success | **yes** |
| Frames, then a zero-length frame, but every frame was a heartbeat | Synthesis produced nothing | no — refunded |
| The body just stops, no zero-length frame | Cut short mid-generation | no — refunded |
| `503` with `"code": "STREAM_BUSY"` | Every node is healthy but streaming capacity is used up | no — nothing was charged |
| `502` | Every node for this engine failed | no — refunded |
`STREAM_BUSY` carries `"fallback": "generate"`, and that is a real instruction:
the queued `POST /v1/tts` path has separate capacity and will accept the work
right now. Retrying the stream immediately usually will not.
A stream that reaches us but produces no audio for **120 seconds** is abandoned
server-side and ends without the zero-length frame — so it arrives as the
"cut short" row above. That ceiling exists so a wedged node cannot hold your
connection open indefinitely.
### If your client disconnects
**You are charged.** Closing the connection part-way — a user pressing Stop, a
timeout on your side, a crashed worker of your own — bills the request. You
received audio; we generated it.
This is deliberate, and it is the one case where "no zero-length frame" does not
mean a refund. If you retry after aborting, budget for both attempts.
### Limits
| Limit | Value | On exceeding |
|---|---|---|
| Concurrent streams | 4 | `429`, `code: STREAM_CONCURRENCY` |
| Synthesis requests per minute | 300 | `429`, honour `Retry-After` |
| Text length | 50 000 characters | `400` |
Concurrency is counted against the **token grant** behind your key, not the key
itself, so several keys issued on one account normally share the four rather than
each getting four of their own. The slot is taken before the balance is touched:
a stream refused here costs nothing.
:::caution The stream route has a much tighter connection cap
`POST /v1/tts/stream` is matched by its own nginx location — the one that turns
response buffering off, without which the whole point of streaming is lost. That
location does not inherit `/api/v1`'s generous per-address connection allowance
of 300; it falls back to the server-wide one of **20 concurrent connections per
source IP**, and to a per-address request rate of 500/min rather than 3000.
A stream holds its connection open for the whole synthesis, so those 20 are held,
not cycled. Twenty concurrent streams from one egress address is a realistic
number for a busy integration, and the twenty-first is refused by nginx with a
bare `503` — an HTML error page, no `Retry-After`, no `X-RateLimit-*`, and none
of the JSON a `STREAM_BUSY` carries. That absence is how you tell the two `503`s
apart: ours has a body and a `"fallback"` you can act on, the edge's has neither.
It is also counted across everyone behind that address, not just you. See
[Rate limits](./overview.md#rate-limits-and-request-ids).
:::
## Cloned voices
Pass a `clone_…` id as `voiceId` and it streams like any other voice — the same
frames, the same terminator, the same headers:
```json
{ "text": "…", "voiceId": "clone_9f2c1e04-…" }
```
The ids come from `GET /v1/voices` with your API key (they are listed alongside
the catalogue, tagged `"kind": "cloned"`). You create the voice itself in the web
Studio at [vieneu.io/#/clone](https://vieneu.io/#/clone) — the API no longer
enrols voices — and it is usable here the moment it is saved. Both your own
clones and any an administrator has published to the catalogue work.
**It costs the same as a preset.** Streaming is billed per submitted character ×
the engine's multiplier, and cloning adds no multiplier of its own. Enrolment is
charged in the Studio, not here, so a stream with a `clone_…` voice costs exactly
what the same text costs with a catalogue voice.
:::caution Clones enrolled before 2026-08-28
Every clone created since 2026-08-28 enrols on `v4`, and those stream exactly as
described above. A clone enrolled earlier was cut for `v3`, and `v3` was retired
from the cloud API on 2026-09-24: such a voice is no longer listed by
`GET /v1/voices` and cannot be rendered through the API on any route — the only
engine that could read its reference clip is the one that now answers `400`.
Re-enrol it in the [Studio](https://vieneu.io/#/clone); the new `clone_…` id
streams like any other `v4` voice. Audio you generated with the old voice before
the retirement is unaffected.
:::
Two failures are worth distinguishing, and both arrive as `400` **before
anything is billed**:
| Message | What happened |
|---|---|
| `Cloned voice "…" was not found among your voices.` | Wrong id, or a clone belonging to another account |
| `Cloned voice "…" is no longer available.` | The voice exists but its reference clip is gone — usually deleted mid-request |
## Reference decoder — Python
```python
import struct, requests
def stream_frames(text, voice, api_key, output_format="wav"):
"""Yield each audio frame. Raises if the stream was truncated."""
resp = requests.post(
"https://api.vieneu.io/api/v1/tts/stream",
headers={"Authorization": f"Bearer {api_key}"},
json={"text": text, "voiceId": voice, "outputFormat": output_format},
stream=True,
)
resp.raise_for_status()
print("sample rate:", resp.headers.get("X-Sample-Rate"))
buf, complete, any_audio = bytearray(), False, False
for chunk in resp.iter_content(chunk_size=8192):
buf.extend(chunk)
# A frame may span chunks, and several may arrive in one.
while len(buf) >= 4:
(length,) = struct.unpack(">I", buf[:4])
if length == 0: # end-of-stream marker
complete = True
del buf[:4]
continue
if len(buf) < 4 + length: # frame not all here yet
break
frame = bytes(buf[4:4 + length])
del buf[:4 + length]
if has_samples(frame): # skip heartbeats — see above
any_audio = True
yield frame
if not complete:
raise RuntimeError("stream truncated — discard this audio")
if not any_audio:
# Kết thúc sạch nhưng không một mẫu nào: chúng tôi cũng không tính tiền
# ca này. Coi nó là thành công là ghi sổ lệch với hoá đơn.
raise RuntimeError("stream carried no audio — not billed, do not save")
```
## Reference decoder — JavaScript
```js
/** False for a heartbeat: a valid WAV whose `data` chunk is empty. */
function hasSamples(frame) {
if (frame.length === 0) return false; // nothing there at all
if (frame.length < 12) return true; // too short to read — assume audio
const dv = new DataView(frame.buffer, frame.byteOffset, frame.byteLength);
const tag = (o) => String.fromCharCode(...frame.subarray(o, o + 4));
if (tag(0) !== 'RIFF' || tag(8) !== 'WAVE') return true; // not WAV — assume audio
let off = 12;
while (off + 8 <= frame.length) {
const size = dv.getUint32(off + 4, true);
if (tag(off) === 'data') return size > 0;
off += 8 + size + (size % 2); // RIFF chunks are word-aligned
}
return true;
}
async function* streamFrames(text, voice, apiKey, outputFormat = 'wav') {
const resp = await fetch('https://api.vieneu.io/api/v1/tts/stream', {
method: 'POST',
headers: {
Authorization: `Bearer ${apiKey}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({ text, voiceId: voice, outputFormat }),
});
if (!resp.ok) throw new Error(`HTTP ${resp.status}`);
const reader = resp.body.getReader();
let buf = new Uint8Array(0);
let complete = false;
let anyAudio = false;
for (;;) {
const { done, value } = await reader.read();
if (done) break;
const next = new Uint8Array(buf.length + value.length);
next.set(buf);
next.set(value, buf.length);
buf = next;
for (;;) {
if (buf.length < 4) break;
const length = new DataView(buf.buffer, buf.byteOffset, 4).getUint32(0);
if (length === 0) { // end-of-stream marker
complete = true;
buf = buf.subarray(4);
continue;
}
if (buf.length < 4 + length) break;
const frame = buf.slice(4, 4 + length);
buf = buf.subarray(4 + length);
if (hasSamples(frame)) { anyAudio = true; yield frame; } // skip heartbeats
}
}
if (!complete) throw new Error('stream truncated — discard this audio');
// Kết thúc sạch nhưng không một mẫu nào: chúng tôi cũng không tính tiền ca
// này. Coi nó là thành công là ghi sổ lệch với hoá đơn.
if (!anyAudio) throw new Error('stream carried no audio — not billed, do not save');
}
```
With the default WAV framing, each yielded frame is a complete file — in a
browser you can feed them straight to `decodeAudioData` and queue the results.
The `hasSamples` guard above is what makes that safe: without it a heartbeat
decodes to a zero-length buffer and quietly joins the queue.
## OpenAI-style streaming
If you are driving this from an OpenAI client, skip the framing entirely:
```python
with client.audio.speech.with_streaming_response.create(
model="tts-1", voice="Ngọc Lan", input="…", response_format="pcm",
extra_body={"stream_format": "audio"},
) as resp:
resp.stream_to_file("speech.pcm") # raw s16le, 24 kHz
```
`stream_format: "audio"` gives you the audio bytes as chunked transfer encoding —
append them and you have the whole thing. `stream_format: "sse"` gives
Server-Sent Events: `speech.audio.delta` events carrying base64 audio, ending in
exactly one `speech.audio.done` or `speech.audio.error`.
As with the native stream, **`speech.audio.done` is the proof the audio is
whole**; if it never arrives, discard what you have — you were not billed.
`stream_format` accepts `pcm` and `ulaw` only, for the reason in the warning
above: the other formats cannot be concatenated into a playable result. For a
complete mp3 or opus file, make an ordinary non-streaming request.
## Latency, honestly
**On the standard plans** — Starter through Hội viên, everything that shares the
pooled lane — first audio lands in roughly **1–2 seconds** end to end. That is
fast enough for read-aloud, dubbing, IVR prompts and assistants that tolerate a
beat before speaking. It is **not** the 200–300 ms class that hard real-time
conversational agents expect, and no amount of client tuning moves it: the wait
is the queue plus the opener chunk, not your connection.
**Enterprise is the exception, by contract rather than by luck.** That tier runs
on a dedicated GPU node with a lane nobody else shares, so the latency ceiling
becomes something we agree on — down to **≤ 250 ms** — instead of whatever the
pool happens to be doing. It is a ceiling written into the deal and sized to
your traffic, not a number you can read off the shared endpoint and expect to
hold. If you are building a real-time agent, talk to us before you design around
the 1–2 second figure above.
---
# Errors
Source: https://docs.vieneu.io/docs/cloud-api/errors
Statuses tell you how bad it was. Most refusals additionally carry a `code`, and
where one exists it is the only part of the body safe to branch on — messages are
prose, some of it Vietnamese, and all of it subject to rewording.
The important thing to know first is that **`code` is not universal**. Most of
the ways a `/v1` call can be refused carry one; a few give you a status and a
sentence. This page says which is which, because a client written on the
assumption that every error has a `code` will silently fall through to its
default branch on the ones that do not.
## Two body shapes
Native endpoints return the platform's shape:
```json
{ "statusCode": 403, "message": "Insufficient tokens", "traceId": "0f7c…" }
```
`POST /v1/audio/speech` returns OpenAI's shape instead, so an OpenAI client can
parse it without an adapter:
```json
{ "error": { "message": "…", "type": "rate_limit_exceeded", "param": null, "code": null } }
```
That is the whole difference: one route, one envelope. Everything else on `/v1` —
including `/v1/vapi/speech`, which serves a third party — uses the first shape.
Every native body repeats the status as `statusCode` and adds `code` and its
extras where the refusal carries one:
```json
{ "statusCode": 429,
"message": "Server is at capacity. Please try again in a few minutes.",
"code": "QUEUE_FULL", "traceId": "0f7c…" }
```
A body that carried a `code` used to arrive without `statusCode` — so gaining a
machine-readable reason cost you a field you had always been able to read. That
gap is closed: the field is on every native body now. The HTTP status line is
still the authoritative copy — read it from the response where you can.
## Refusals that carry a `code`
### Quota and grant
The six routes that bill per character or per second — `POST /v1/tts`,
`/v1/tts/stream`, `/v1/dialogue`, `/v1/dub`, `/v1/srt` and
`/v1/vapi/speech` — answer a quota refusal with a code, and the two that recover
on their own say when:
| `code` | Status | What it means | What to do |
|---|---|---|---|
| `GRANT_EXPIRED` | 403 | The token package behind the key has passed its expiry date. | Renew. **Retrying will not help**, and neither will topping up. |
| `GRANT_TOKENS_EXHAUSTED` | 403 | The grant's balance is smaller than this request's price. Carries `remaining` and `required`. | Top up. Retrying will not help. |
| `GRANT_DAILY_LIMIT` | 429 | The plan's daily cap is spent. Carries `resetAt`, `remaining`, `required`. | Sleep until `resetAt`. |
| `GRANT_WEEKLY_LIMIT` | 429 | The plan's weekly cap is spent. Carries `resetAt`, `remaining`, `required`. | Sleep until `resetAt`. |
| `CONCURRENT_CONFLICT` | 429 | Two of your own requests raced for the same balance and this one lost. Carries `retryAfterSeconds: 1`, also sent as `Retry-After: 1`. | Wait the second it names, then retry — this one clears. |
`resetAt` is ISO 8601, and it replaces parsing the timestamp out of the message.
Take it over a guessed backoff, and over an "upgrade your plan" prompt for
something that resets in an hour.
The 403/429 split is the coarse version of the same decision: **403 means
retrying will not help** even though it looks transient, and 429 means it will,
eventually. What the status cannot always tell you is *when*. The daily and
weekly caps carry **no `Retry-After`** — they carry `resetAt` instead, which is
the honest answer for a wait measured in hours. The refusals that clear in
seconds do carry it: `CONCURRENT_CONFLICT` (`Retry-After: 1`) and the capacity
codes below (`Retry-After: 15`). `code` is what separates a race you retry from a
cap you wait out, and `resetAt` is what says how long the wait is.
`FREE_DAILY_LIMIT` and `GRANT_INACTIVE` appear in the platform's internal code
list but not here: an API key always bills a grant, so the free tier is
unreachable, and a grant that is not active answers 403 with prose and no code.
### Capacity
| `code` | Status | What it means | What to do |
|---|---|---|---|
| `QUEUE_FULL` | 429 | The asynchronous job queue is at its depth limit. Carries `retryAfterSeconds` (`Retry-After`). | Wait what the header says and resubmit — with the same `Idempotency-Key` if you sent one. Nothing was charged. |
| `USER_QUEUE_FULL` | 429 | *Your* account already holds its ceiling of queued jobs (30 by default). Carries `retryAfterSeconds`. | Poll what you have queued, then resubmit. Nothing was charged. |
| `STREAM_CONCURRENCY` | 429 | You already hold the maximum number of open streams. | Close one, or wait for one to finish. Nothing was charged. |
| `STREAM_BUSY` | 503 | Every node is healthy but streaming capacity is used up. | Carries `"fallback": "generate"`, and that is a real instruction — the queued path has separate capacity. See [How a stream ends](./streaming.md#how-a-stream-ends). |
`QUEUE_FULL` and `USER_QUEUE_FULL` reach you from `POST /v1/tts` only; the
synchronous routes do not queue.
### Idempotency
Only on `POST /v1/tts`, and only when you sent an `Idempotency-Key` — see
[Retrying safely](./overview.md#retrying-safely).
| `code` | Status | What it means | What to do |
|---|---|---|---|
| `IDEMPOTENCY_KEY_INVALID` | 400 | The header is not 8–128 printable ASCII characters without whitespace. | Send a UUID. |
| `IDEMPOTENCY_IN_FLIGHT` | 409 | A request with this key is still being processed. Carries `retryAfterSeconds: 1` (`Retry-After: 1`). | Retry after the second, with the **same** key. A new key can charge you twice. |
| `IDEMPOTENCY_KEY_REUSED` | 422 | This key was already used with a different request body. | Use a fresh key for a different request; keys are per request, not per session. |
### Cloning
Cloning moved to the web Studio. The enrolment routes — `POST /v1/voices`,
`DELETE /v1/voices/{voiceId}`, `POST /v1/clone`, `POST /v1/upload` and
`POST /v1/prepare` — answer one refusal now, before any billing or GPU work:
| `code` | Status | What it means | What to do |
|---|---|---|---|
| `CLONE_WEB_ONLY` | 410 | Cloned voices are created at [vieneu.io/#/clone](https://vieneu.io/#/clone), not through the API. Carries `cloneUrl`. | Make the voice in the Studio once. Nothing else in your integration changes: it is then listed by `GET /v1/voices` with your key and its `clone_…` id works as `voiceId` on `POST /v1/tts` and `/v1/tts/stream`. |
410 rather than 404 on purpose: the endpoints existed and were documented, so
the status says the surface moved rather than that you mistyped a path. The
older clone codes (`CLONE_QUOTA_EXCEEDED`, `CLONE_MONTHLY_CAP`,
`CLONE_DAILY_CAP`, `CLONE_REF_*`, `CLONE_TRANSCRIPT_DENSITY`,
`CLONE_TOKENS_INSUFFICIENT` and the rest) belong to those routes and can no
longer reach an API key — the same limits still apply to the Studio, where they
are shown as you clone.
Generating **with** a cloned voice is not part of this and carries no clone
codes. A `clone_…` id that is missing or belongs to someone else fails as an
ordinary `400` with a message and no `code` — see
[Cloned voices](./streaming.md#cloned-voices).
## Refusals that do **not** carry a code
### The OpenAI-compatible route
`POST /v1/audio/speech` carries none of the codes above, despite OpenAI's
envelope having a field for it: its `code` is always `null` on this route. What
you get instead is `type`, which is coarser on purpose. A quota refusal is typed
from its status — `insufficient_quota` on the `403`s, which are the exhausted
and expired grants, and `rate_limit_exceeded` on the `429`s, which are the spent
caps and the races. So "out of credit" and "wait" survive the translation;
`GRANT_EXPIRED` and `GRANT_TOKENS_EXHAUSTED` do not, and neither does the
`resetAt` that would tell you how long to wait. Anything raised before the
handler runs — a bad key, a field the validator rejects — is typed from the
status alone, and that mapping folds `401` and `403` together as
`authentication_error`.
### The public API never returns 402
There is no `402 Payment Required` anywhere on `/v1`. Every money path runs
through the same deduction, and that step answers only `403` or `429`. The spec
has carried a `402` example; the server does not send one. A client that treats
`402` as "out of credit" will read the real `403` as a permissions bug and stop
retrying for the wrong reason.
### Two text validations
Before any tokens are spent, submitted text is checked for two things the model
cannot usefully voice. Both come back as a plain `400` with a message and no
`code`.
| Message begins | Rule |
|---|---|
| `Text appears to be in an unsupported language…` | More than **34%** of the letters are outside the Latin script. Vietnamese diacritics and `đ` are Latin, so this fires on Chinese, Japanese, Korean, Cyrillic and the like — including an otherwise Vietnamese passage carrying a long non-Latin quotation. |
| `Text does not look like readable words…` | Either one unbroken run of **more than 30 letters**, or **more than 60%** of the word-like tokens (four letters or longer, and at least three of them present) carry no vowel. |
Punctuation breaks a letter run, so a URL, an email address or a file path
measures as its parts and passes; a keyboard mash still measures as one run and
does not.
**These fire on three routes only:** `POST /v1/tts`, `POST /v1/dialogue` and
`POST /v1/srt`. `POST /v1/tts/stream`, `POST /v1/audio/speech`,
`POST /v1/vapi/speech` and `POST /v1/dub` do not run the check — so text one
route rejects, another will synthesize and bill you for. Worth knowing before you
treat a `400` from one route as proof the text is bad everywhere.
## Stream concurrency
`POST /v1/tts/stream` refuses a new stream while you already hold too many:
| Counted against | Ceiling | On exceeding |
|---|---|---|
| The token grant behind your API key | 4 | `429`, `code: STREAM_CONCURRENCY` |
| A signed-in user in the web app | 2 | `429`, same code |
The slot is taken **before** the balance is touched, so a refused stream leaves
no mark on your tokens. Two details about the counting: it is per backend
process, and the key is the **token grant**, not the API key — so several keys
issued against one grant share the four rather than each getting four of their
own.
## `X-Request-Id`
Every response carries `X-Request-Id`. It is the id our logs are keyed by: one
value spans the API request, the worker call it made, and every log line either
produced.
Most error bodies repeat the value as `traceId`, but **the header is the copy to
read** — it is the only one that is always there. `traceId` is stamped by the
exception filter, so a body written straight to the response never gets one:
- **`POST /v1/audio/speech`**, on every error. The handler writes OpenAI's
envelope itself, and the filter that would add `traceId` never runs — nor does
OpenAI's shape have a field for it.
- **`POST /v1/vapi/speech`**, on the `502` it sends when the worker produced no
audio. That body is written mid-response and carries `statusCode` and
`message` only.
Quote it in a support request. Without it, "a request failed around 3pm" is a
search; with it, it is a lookup.
The value is always ours. The edge proxy sets `X-Request-ID` on every request it
forwards, unconditionally, so an id you send on the way in is overwritten rather
than adopted — log the one that comes back next to your own correlation id
instead of expecting yours to survive.
## Status summary
| Status | Meaning |
|---|---|
| 400 | Malformed request — an unknown voice, a bad format, text the validator refused, a clone rejection |
| 401 | Missing, malformed or revoked API key |
| 403 | Out of tokens, grant expired, or your plan does not include this engine or feature |
| 413 | Uploaded file over 10 MB |
| 409 | An `Idempotency-Key` whose first call is still running — retry with the same key |
| 422 | Content refused by moderation (only when `aiRefine` is on), or an `Idempotency-Key` reused with a different body |
| 429 | Rate limit, token quota, queue depth or stream concurrency |
| 500 | Synthesis failed. The charge is refunded automatically |
| 502 | Every worker for this engine failed. Refunded |
| 503 | No worker for the requested engine or format, or streaming capacity is full — retry shortly |
`429` is the one status with unrelated causes behind it, and the headers say
which — see [Rate limits](./overview.md#rate-limits-and-request-ids).
---
# Webhooks
Source: https://docs.vieneu.io/docs/cloud-api/webhooks
`POST /v1/tts` is asynchronous. Instead of polling `GET /v1/tts/{jobId}` until the
status changes, register an HTTPS endpoint and we will POST a signed event to it
the moment the job reaches a terminal state.
Polling still works and is still the recovery path when a delivery fails — see
[When we give up](#when-we-give-up).
## Events
| Type | When |
| --- | --- |
| `tts.job.completed` | The job produced audio. The event carries a short-lived download URL. |
| `tts.job.failed` | The job failed permanently after all retries. Tokens have been refunded. |
Only `POST /v1/tts` produces jobs today, so those are the only two event types.
Every other synthesis route (`/v1/audio/speech`, `/v1/tts/stream`, `/v1/dialogue`,
`/v1/dub`, `/v1/srt`) answers inline and produces no event.
## Register an endpoint
```bash
curl -X POST https://api.vieneu.io/api/v1/webhooks \
-H "Authorization: Bearer vn_sk_…" \
-H "Content-Type: application/json" \
-d '{
"url": "https://api.example.com/hooks/vieneu",
"description": "Production receiver"
}'
```
```json
{
"id": "665f1a2b3c4d5e6f7a8b9c0d",
"url": "https://api.example.com/hooks/vieneu",
"events": ["tts.job.completed", "tts.job.failed"],
"status": "ENABLED",
"secretPrefix": "whsec_a1b2c3",
"secret": "whsec_2f9c…"
}
```
**`secret` is returned once and never again.** Store it now. If you lose it, call
`POST /v1/webhooks/{id}/rotate-secret` and update your receiver.
### Destination requirements
- **`https` only.** There is no `http` option, in any mode. Events carry a
presigned download URL for your audio, which is a bearer credential — the
signature protects integrity, not confidentiality. For local development use a
tunnel (ngrok, Cloudflare Tunnel, the like) rather than an `http` URL.
- **Public addresses only.** Private, loopback, link-local, CGNAT and reserved
ranges are refused — when you register the endpoint, and again at the moment of
every delivery. A hostname that resolves publicly at registration and privately
later is refused at connect time.
- **No redirects.** A `3xx` response is treated as a failed delivery. Point the
endpoint at the final URL.
- **No credentials in the URL.** `https://user:pass@…` is refused.
### Scoping to one API key
Pass `apiKeyId` when registering to receive only the jobs submitted with that
key. Omit it and the endpoint receives events for every key on the account.
Useful for keeping a test key's traffic off your production receiver.
## The event body
```json
{
"id": "evt_9f2c1e043b8a4d218f770c1a2b3c4d5e",
"type": "tts.job.completed",
"api_version": "2026-09-01",
"created": 1756684800,
"attempt": 1,
"data": {
"job_id": "0f7c1a2b-3c4d-5e6f-7a8b-9c0d1e2f3a4b",
"status": "completed",
"voice_id": "Ngọc Lan",
"engine": "v4",
"character_count": 412,
"token_cost": 1236,
"duration_seconds": 27.4,
"processing_time_ms": 8120,
"audio_url": "https://…s3…?X-Amz-Signature=…",
"audio_url_expires_at": "2026-09-01T11:00:00.000Z"
}
}
```
A failure looks like this:
```json
{
"id": "evt_1a4d…",
"type": "tts.job.failed",
"api_version": "2026-09-01",
"created": 1756684800,
"attempt": 1,
"data": {
"job_id": "0f7c1a2b-3c4d-5e6f-7a8b-9c0d1e2f3a4b",
"status": "failed",
"voice_id": "Ngọc Lan",
"engine": "v4",
"character_count": 412,
"token_cost": 1236,
"error": {
"code": "synthesis_unavailable",
"message": "No synthesis worker could complete the job. Tokens were refunded."
},
"refunded": true
}
}
```
Branch on `error.code`, never on `error.message`. The codes are
`synthesis_failed`, `synthesis_unavailable`, `synthesis_timeout`,
`content_rejected` and `quota_exceeded`; the message is prose and may be
reworded.
### About `audio_url`
`audio_url` is minted fresh for **each delivery attempt** and expires one hour
after that attempt — `audio_url_expires_at` tells you exactly when. It is
deliberately shorter-lived than the 24-hour URL `GET /v1/tts/{jobId}` returns,
because an event body may end up in your logs. Download promptly, or call
`GET /v1/tts/{jobId}` for a fresh URL whenever you need one.
The event carries everything you need to act on the job without a second API
call. It does **not** carry the submitted text.
## Verifying the signature
Every request carries a `Vieneu-Signature` header:
```
Vieneu-Signature: t=1756684800,v1=5257a869e7ecebeda32affa62cdca3fa51cad7e77a0e56ff536d0ce8e108d8bd
```
- `t` is the unix timestamp of **this delivery attempt**.
- `v1` is `HMAC-SHA256(secret, "." + rawBody)`, hex-encoded.
During a secret rotation the header carries **one `v1=` element per live
secret**. Check the signature against each of them and accept if any matches.
Four rules that matter:
1. **Use the raw request body bytes.** Not a re-serialised object. `JSON.parse`
then `JSON.stringify` changes key order and unicode escaping, and the
signature will fail intermittently in a way that is very hard to debug. Read
the raw body before any JSON middleware touches it.
2. **Reject a stale `t`.** Use a tolerance of **300 seconds**. Do not use `0` —
that disables the recency check entirely.
3. **Ignore any scheme that is not `v1`.** If a future element `v2=` appears,
a verifier that accepts "any element that matches" can be downgraded.
4. **Compare in constant time.** `crypto.timingSafeEqual`, not `===`.
### Node.js
```js
const crypto = require('crypto');
/**
* Verify a Vieneu webhook signature.
*
* @param {Buffer|string} payload RAW request body — not a parsed object.
* @param {string} header The `Vieneu-Signature` header value.
* @param {string} secret Your `whsec_…` signing secret.
* @param {number} toleranceSeconds Max age of `t`. 300 is the documented value.
* @returns {boolean}
*/
function verifySignature(payload, header, secret, toleranceSeconds = 300) {
const body = Buffer.isBuffer(payload) ? payload : Buffer.from(payload, 'utf8');
let timestamp = null;
const signatures = [];
for (const element of String(header || '').split(',')) {
const idx = element.indexOf('=');
if (idx === -1) continue;
const key = element.slice(0, idx).trim();
const value = element.slice(idx + 1).trim();
if (key === 't') timestamp = Number(value);
// Only v1. Ignoring unknown schemes is what stops a downgrade.
else if (key === 'v1') signatures.push(value);
}
if (timestamp === null || !Number.isFinite(timestamp)) return false;
if (signatures.length === 0) return false;
// Replay window. A tolerance of 0 disables this check — don't.
if (toleranceSeconds > 0) {
const now = Math.floor(Date.now() / 1000);
if (Math.abs(now - timestamp) > toleranceSeconds) return false;
}
const expected = crypto
.createHmac('sha256', secret)
.update(Buffer.concat([Buffer.from(timestamp + '.', 'utf8'), body]))
.digest('hex');
return signatures.some((candidate) => {
// timingSafeEqual throws on a length mismatch, so check length first —
// the length of a hex digest is not itself a secret.
if (candidate.length !== expected.length) return false;
try {
return crypto.timingSafeEqual(
Buffer.from(candidate, 'hex'),
Buffer.from(expected, 'hex'),
);
} catch (err) {
return false;
}
});
}
```
Wire it up in Express, taking care to keep the raw body:
```js
const express = require('express');
const app = express();
app.post(
'/hooks/vieneu',
express.raw({ type: 'application/json' }),
(req, res) => {
const ok = verifySignature(
req.body, // Buffer, thanks to express.raw
req.get('Vieneu-Signature'),
process.env.VIENEU_WEBHOOK_SECRET,
);
if (!ok) return res.sendStatus(400);
// Answer FIRST, work afterwards. We time out at 10 seconds.
res.sendStatus(200);
const event = JSON.parse(req.body.toString('utf8'));
if (alreadyProcessed(event.id)) return; // at-least-once — dedupe on id
void handle(event);
},
);
```
### Python
```python
import hashlib
import hmac
import time
def verify_signature(payload: bytes, header: str, secret: str, tolerance_seconds: int = 300) -> bool:
timestamp = None
signatures = []
for element in (header or "").split(","):
key, sep, value = element.partition("=")
if not sep:
continue
key, value = key.strip(), value.strip()
if key == "t":
try:
timestamp = int(value)
except ValueError:
return False
elif key == "v1": # only v1; ignore any other scheme
signatures.append(value)
if timestamp is None or not signatures:
return False
if tolerance_seconds > 0 and abs(int(time.time()) - timestamp) > tolerance_seconds:
return False
expected = hmac.new(
secret.encode("utf-8"),
f"{timestamp}.".encode("utf-8") + payload,
hashlib.sha256,
).hexdigest()
return any(hmac.compare_digest(candidate, expected) for candidate in signatures)
```
## Delivery guarantees
**At-least-once, unordered.** Concretely:
- **Duplicates are normal.** Design your receiver to be idempotent. Dedupe on
`id`, which is stable for a given job and event type — the same terminal state
re-emitted after a crash or a queue redelivery carries the same `id`.
- **`t` and `v1` change between duplicates.** Every attempt is signed afresh.
Never dedupe on the signature.
- **Do not use `created` for ordering or deduplication.** It is a diagnostic
timestamp. Two jobs' events can arrive in either order, and a retry of an
older event can land after a newer one.
- **`attempt` tells a retry from a first delivery.** It is 1-based, and the
`Vieneu-Delivery` header identifies one specific HTTP attempt.
### Other headers
| Header | Meaning |
| --- | --- |
| `Vieneu-Signature` | `t=…,v1=…` — see above. |
| `Vieneu-Event-Id` | Same value as `id` in the body. Convenient for dedupe at the edge. |
| `Vieneu-Event-Type` | Same value as `type` in the body. |
| `Vieneu-Delivery` | Identifies ONE http attempt. Not a dedupe key. |
## What we expect from your endpoint
- **Answer `2xx`.** Anything else — including any `3xx` — is a failed delivery.
- **Answer within 10 seconds.** Acknowledge first, do the work after.
- **Keep the response small.** We stop reading as soon as 64 KB has arrived (we
finish the chunk in flight, so a little more may cross the wire) and discard it. We never
log, store or return your response body.
## Retries and backoff
A failed delivery is retried up to **6 attempts**, with the gaps growing each
time. Measured against the queue library we actually run, they fire at:
```
t+0 t+0.5m t+2.0m t+5.5m t+13.0m t+28.5m
```
So the window is about **28.5 minutes** end to end, and the longest single gap
between two attempts — the one before the last — is **15.5 minutes**.
That second number is the one to build against. **Size any staleness or
reconciliation threshold above 15.5 minutes.** A sweeper that treats a delivery
as stranded after fifteen will keep finding deliveries that are simply waiting
out their final backoff, and will re-enqueue work that was never lost.
:::note This page used to say 15 minutes
It described the schedule as "+30s, +1m, +2m, +4m, +8m, about 15 minutes". That
was arithmetic on the wrong formula — the queue computes each delay as
`(2^attempts − 1) × 30s`, not `30s × 2^attempts`, which roughly doubles both the
window and every gap inside it. The numbers above were read off the installed
library rather than recalled.
:::
## When we give up
After the last attempt the event is **dead-lettered**: it will not be sent
again. Two things happen:
1. The delivery row is marked `DEAD`, visible at
`GET /v1/webhooks/{id}/deliveries`.
2. A notification appears in your Vieneu account.
The audio is not lost. `GET /v1/tts/{jobId}` still returns the job and a fresh
24-hour download URL. If your receiver was down, reconcile from there.
## Self-diagnosis
```bash
curl https://api.vieneu.io/api/v1/webhooks/{id}/deliveries \
-H "Authorization: Bearer vn_sk_…"
```
```json
[
{
"id": "665f…",
"eventId": "evt_9f2c…",
"eventType": "tts.job.completed",
"status": "DEAD",
"attempts": 6,
"lastStatusCode": 502,
"lastError": "http_502",
"lastDurationMs": 143,
"lastAttemptAt": "2026-09-01T10:15:02Z",
"deliveredAt": null
}
]
```
`lastError` values you may see:
| Code | Meaning |
| --- | --- |
| `http_4xx` / `http_5xx` | Your endpoint answered with that status. |
| `redirect_refused` | Your endpoint returned a `3xx`. Point it at the final URL. |
| `connect_timeout` / `response_timeout` | Your endpoint did not answer in time. |
| `connection_refused` / `dns_failure` / `tls_failure` | We could not reach it. |
| `blocked_private_address` | The URL resolves to a non-public address. |
| `endpoint_unavailable` | The endpoint was deleted or disabled mid-flight. |
| `audio_not_uploaded` | The audio had not finished uploading yet. Retried. |
## Rotating the signing secret
```bash
curl -X POST https://api.vieneu.io/api/v1/webhooks/{id}/rotate-secret \
-H "Authorization: Bearer vn_sk_…"
```
The new secret is returned once. The previous secret keeps verifying for **24
hours**, and during that window each event carries one `v1=` per live secret —
so you can deploy the new secret without dropping events signed with the old
one. The verifier above already handles this: it accepts if any `v1` matches.
---
# Vapi (voice agent)
Source: https://docs.vieneu.io/docs/cloud-api/vapi
[Vapi](https://vapi.ai) builds voice agents that answer and place phone calls. It
speaks through a TTS provider of your choosing — and its `custom-voice` provider
lets that be any HTTPS endpoint. VieNeu implements exactly what it expects, so
your agent can answer in Vietnamese.
This matters because Vietnamese is essentially absent from the realtime TTS
vendors Vapi ships with.
## Configure the assistant
```json
{
"voice": {
"provider": "custom-voice",
"server": {
"url": "https://api.vieneu.io/api/v1/vapi/speech?voiceId=Ngọc%20Lan",
"secret": "vn_sk_your_key_here",
"timeoutSeconds": 30
}
}
}
```
That is the whole integration. Two details are doing the work:
**`secret` is your VieNeu API key.** Vapi sends it as the `X-VAPI-SECRET` header,
and its config has no field for an `Authorization` header — so this is where the
key goes. It is the same key, checked the same way, and it is billed to the same
account.
**The voice goes in the URL.** Vapi's request payload has no voice field, so pass
`?voiceId=` (URL-encoded). List the options with `GET /v1/audio/voices`. Add
`&engine=v4` for the premium engine if your plan includes it. Omit `voiceId`
entirely and you get the engine's default voice.
## What happens on each utterance
Vapi POSTs:
```json
{
"message": {
"type": "voice-request",
"text": "Xin chào, tôi có thể giúp gì cho bạn?",
"sampleRate": 24000
}
}
```
VieNeu answers `200` with `Content-Type: application/octet-stream` and raw mono
16-bit little-endian PCM at exactly that sample rate, streamed as it is
generated rather than buffered — so the agent starts speaking sooner.
Vapi asks for 8000, 16000, 22050 or 24000 Hz depending on the transport. All four
are supported, and the response is always at the rate requested: raw PCM carries
no rate of its own, so anything else would come out at the wrong pitch.
## Billing and behaviour
Billed per submitted character, like the rest of `/v1`. Audio that gets cut short
is refunded automatically.
**AI refinement never runs on this endpoint**, regardless of your account
settings. An agent's job is to answer quickly, and an extra model round-trip in
front of every utterance is the opposite of that. Deterministic text preparation
still runs, so Vietnamese is pronounced correctly. If your agent's text contains
things you want read a particular way — currency, dates, product codes — normalize
it in your prompt, where you can see the result.
## Latency, honestly
**On the standard plans**, first audio leaves in roughly 1–2 seconds. That works
for an agent that answers a question, reads a menu, confirms a booking. It is
**slower than the 200–300 ms** the fastest English-only vendors reach, and it
will be noticeable as a beat before the agent speaks.
An Enterprise deal runs on a dedicated node where the latency ceiling is agreed
per contract — down to ≤ 250 ms — which is the version of this you want if the
beat is a dealbreaker. See
[Latency, honestly](./streaming#latency-honestly).
Two things worth doing: keep `timeoutSeconds` at 30 or above, and keep utterances
short — a paragraph costs more waiting than three sentences.
**Utterances are capped at 800 characters** (about 50 seconds of speech), and the
cap exists for that reason: a longer turn cannot be delivered inside Vapi's
30-second deadline, so it would be cut off mid-word and billed anyway. If your
agent produces longer replies, split them — Vapi will request each piece
separately and the caller hears them back to back.
If no worker starts answering within 12 seconds, the request fails with a `503`
rather than holding the socket until Vapi gives up. That distinction matters:
a status code your `fallbackPlan` can act on beats a timeout, which just looks
like a provider that stopped responding.
Set a `fallbackPlan` on the assistant if a missed utterance would be worse than a
non-Vietnamese voice.
## Troubleshooting
| Symptom | Cause |
|---|---|
| `401` | The `secret` is not a valid VieNeu API key, or the key was revoked. |
| `400` with a voice message | `voiceId` is not in the catalogue for that engine — check `GET /v1/audio/voices`. |
| `403` | The key's token grant is empty or expired, or your plan excludes the engine. |
| `400` about `message.text` length | The turn was over 800 characters — split it. |
| `503` naming a sample rate | No worker in the pool has been updated to encode raw PCM yet. We fail over across the fleet first, so this means all of them. Retry; contact us if it persists. |
| `503` about no worker answering | Nothing started producing audio within 12 seconds. Usually a capacity spike; your `fallbackPlan` covers the turn. |
| `502` about no audio | The worker accepted the request and then produced nothing. Not billed. |
Every response carries `X-Request-Id`. Quote it and we can find the exact call.
---
# Changelog
Source: https://docs.vieneu.io/docs/cloud-api/changelog
Changes to the Cloud API (`/api/v1`). Anything marked **⚠️ Behaviour change** can
affect a running integration of yours even if you change nothing.
---
## 2026-10-02 — Long audiobooks: nothing lost, nothing stuck {#2026-10-02--audiobooks-resilience}
### ⚠️ Behaviour change
**A chapter longer than the plan can render is refused when added.** A chapter
is billed in one go, so on a plan with a daily or weekly token cap, a chapter
costing more than the cap could never start: it waited for tokens forever.
`POST /v1/audiobooks/{bookId}/chapters` now answers 400 `chapter_over_plan_limit`
with `chapterCharsMax`, and every book reports `chapterCharsMax`. Plans without
such a cap are unchanged (60,000 characters).
### New
**`afterSegment` when continuing a chapter.** With `appendTo`, pass how many
segments the chapter has; a different count answers 409 `segment_mismatch` with
the words the chapter ends with, so a retried request cannot add its text twice
and a lost one cannot leave a gap. Without it, a request repeating the chapter's
last segments answers 409 `duplicate_segments`, and a new chapter identical to
the last one 409 `duplicate_chapter`. The MCP tool `add_audiobook_chapter` takes
`after_segment` and shows how each chapter now ends. See
[Audiobooks](../integrations/mcp/audiobooks.md#limits).
### Changed
**Failed chapters are made again on their own.** A chapter whose job fails — a
deploy replacing the server twice during one long chapter, a voice server down —
is refunded as before, then rendered again a few minutes later, at most twice,
reusing the parts already made. A job that stops moving for 45 minutes is failed
and refunded the same way, so it can no longer hold up every other book.
Mastering to MP3 retries with growing waits for about an hour before handing over
the WAV, and a take cut short in transit, or an MP3 shorter than its take, is
refused instead of delivered.
**MP3 peaks stay under −3 dBTP.** Mastering now aims half a decibel lower and
adds a peak limiter before the MP3 encoder, which lifts peaks a little: two
chapters of a test book rendered on production had reached −2.8 dBTP. Loudness
is unchanged at −19 LUFS.
---
## 2026-10-01 — Audiobooks, over the API and in the MCP server {#2026-10-01--audiobooks}
### New
**Audiobooks.** A book is a narrator and up to 30 character voices reading
chapters written as a script — narration, and each line of dialogue with the
character who says it. `POST /v1/audiobooks` creates one,
`POST /v1/audiobooks/{bookId}/chapters` adds chapters, `POST /v1/audiobooks/{bookId}/start`
queues it and `GET /v1/audiobooks/{bookId}` follows it, with a download link for
every finished chapter. Books are made in the background, one chapter at a time,
while the GPU fleet is idle, so they take hours rather than seconds. Each chapter
is billed per character when it starts, refunded if it fails, mastered to MP3
(−19 LUFS, peaks ≤ −3 dBTP, silence at both ends, ID3 tags) and saved to the
Library. LIVE keys and catalogue voices only. See
[Every operation](./overview.md#every-operation).
**Audiobooks in the MCP server.** Six tools and a chapter player let Claude or
ChatGPT turn a text the user provides into a multi-voice audiobook. See
[Audiobooks](../integrations/mcp/audiobooks.md).
**"What can VieNeu do?" in the MCP server.** A `list_capabilities` tool (a menu
card in Claude and ChatGPT: click an example to send it) and five ready-made
prompts. Optional tool arguments now also accept `null`. See
[MCP server](../integrations/mcp/index.md#prompts).
**Connected before this release? Reconnect.** Apps keep the tool list they
fetched when they connected, so the new tools appear after you disconnect
VieNeu and connect again. From now on the server's `serverInfo.version`
changes whenever its tools do (`1.2.0+`), for apps that refresh on
it. See [Troubleshooting](../integrations/mcp/troubleshooting.md).
**Voice search that reads a casting brief.** In the MCP server, `list_voices`
takes `search` the way assistants write it, "giọng nam trầm ấm miền Bắc":
gender and region come from the voice's fields, a word written with diacritics
must match them, words such as "giọng" are ignored, and when no voice has every
word the closest come back with what each lacks. See
[MCP server](../integrations/mcp/index.md#list_voices).
---
## 2026-10-01 — MP3 and an inline player in the MCP server; `X-Vieneu-Job-Id` on `/v1/audio/speech` {#2026-10-01--mcp-mp3-and-inline-player}
### New
**The MCP server returns MP3 and plays it in the chat.** Texts up to 1,500
characters now come back as MP3 (about a tenth of the WAV size) instead of WAV,
and in Claude the result shows an audio player right in the conversation. See
[MCP server](../integrations/mcp/index.md#the-inline-player).
**`POST /v1/audio/speech` names its saved copy.** A successful non-streaming
response now carries `X-Vieneu-Job-Id`: the id under which VieNeu keeps a copy
of the audio you just received. A moment later `GET /v1/tts/{jobId}` answers
`completed` with a presigned `audioUrl` for the same bytes — useful when you
want a shareable link rather than the raw file. The header is exposed to
browsers through CORS.
**`GET /v1/balance`.** How many tokens the calling key can spend right now,
the plan it bills, and that plan's daily and weekly caps with their reset
times. Not billed. The MCP server exposes the same thing as the
`get_token_balance` tool and adds the remaining balance to every finished
synthesis.
**MCP sign-in works with ChatGPT.** The authorization server now accepts a
client metadata document that prefers `private_key_jwt` as long as it also
allows `none` (ChatGPT's does), and dynamic registration accepts
`client_secret_post` / `client_secret_basic`.
---
## 2026-09-30 — Hosted MCP server: VieNeu inside Claude and ChatGPT {#2026-09-30--hosted-mcp-server}
### New
**`https://api.vieneu.io/mcp`** — add this URL to Claude, ChatGPT, Cursor or
Claude Code, sign in with your VieNeu account, and ask for Vietnamese speech in
plain language. The assistant searches voices, synthesizes and returns a
download link. No package to install and no API key to paste: the assistant
signs in through OAuth 2.1 and you approve it on a VieNeu page. Tool calls are
billed exactly like the equivalent `/v1` calls. See [MCP server](../integrations/mcp/index.md) and the per-app guides for
[Claude](../integrations/mcp/claude-ai.md), [ChatGPT](../integrations/mcp/chatgpt.md) and others.
Connected apps are listed, and can be disconnected, on the Developer page.
---
## 2026-09-25 — A cache hit on `POST /v1/tts` answers `completed`
### New
**`POST /v1/tts` returns the audio URL at once when the audio already exists.**
Identical requests — the same text, voice, speed, emotion and engine, from any
account — are served from a content cache rather than synthesized again. Such a
call used to answer `status: "queued"` like every other, sending you off to poll
for audio that was already there: one full poll interval spent on nothing.
It now answers `status: "completed"` with `audioUrl`, `audioUrlExpiresIn` and
`duration` — exactly the body `GET /v1/tts/{jobId}` would give — so there is
nothing to poll. Billing is unchanged (a cache hit is charged, as before).
If your client always polls after a submit, nothing breaks: the poll returns the
same `completed` body. To take the shortcut, branch on `status` in the submit
response:
```js
const job = await submit(text, voiceId);
const done = job.status === 'completed' ? job : await pollUntilDone(job.jobId);
```
A submit that is *not* a cache hit still answers `queued` — and so does an
`Idempotency-Key` replay of an earlier submit, whatever that job's state; poll it.
The cached row's bytes are occasionally still on their way to storage, in which
case the response stays `queued` and the poll reports `completed` a moment later.
---
## 2026-09-24 — `v3` retired from the cloud API
### ⚠️ Behaviour change
**The `v3` engine is gone from `/api/v1`.** From 2026-09-24 08:00 (GMT+7) the
cloud API renders on `v4` only. Any request that names `"engine": "v3"` — on
`POST /v1/tts`, `/v1/tts/stream`, `/v1/audio/speech`, `/v1/dialogue`,
`/v1/dub`, `/v1/srt` and their multipart variants — is refused with `400`
before anything is billed, and the message says so:
```
Engine "v3" was retired on 2026-09-24. Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues share no ids.
```
A request that names no engine was already landing on `v4`, the default, and
is not affected — unless it carries a `v3` voice id, which is the next point.
**What the catalogue endpoints say now.** `GET /v1/engines` lists a single
engine, `{ "key": "v4", "isDefault": true, "sampleRate": 48000,
"billingMultiplier": 3, "features": ["generate", "clone", "dialogue", "stream",
"dub", "srt"] }`, with `"globalMultiplier": 1.3` beside it. `GET /v1/voices`
lists `v4` voices only, plus your own `v4` clones when you send a key; `v3`
voices, and clones enrolled on `v3`, are no longer listed. `GET
/v1/voices?engine=v3` is a `400` with the message above. The retired
catalogue's ids — mostly opaque `vieneu-…` slugs — shared almost no space with
`v4`'s display names, so a hard-coded `v3` voice keeps failing on `voiceId` even
once you drop `engine`: take a fresh id from `GET /v1/voices?engine=v4`.
**On the OpenAI-compatible route**, `"model": "vieneu-v3"` answers `400` with
`Model 'vieneu-v3' is retired. Engine "v3" was retired on 2026-09-24. …`.
`tts-1`, `tts-1-hd`, `gpt-4o-mini-tts`, `vieneu` and `vieneu-v4` all resolve to
`v4`. The `emotion` field (`natural` / `storytelling`) was `v3`-only; it is
still accepted, and ignored — `v4` has no reading styles.
**`GET /v1/emotion-tags` with no `engine`** now answers for the default engine,
`v4`: no styles and three inline cues, `[cười]`, `[thở dài]` and
`[hắng giọng]`. It used to answer with `v3`'s set. `?engine=v3` is a `400`. The
`/v1/audio/…` equivalents behave the same way.
**Clones enrolled on `v3`** — anything created before 2026-08-28, when
enrolment moved to `v4` — can no longer be rendered through the API, because
the only engine that could read their reference clip is the one that now
answers `400`. Naming one as `voiceId` (or as the `voice` of `/v1/audio/speech`)
answers `400` with `code: "CLONE_ENGINE_RETIRED"` and the voice's `engine`,
instead of quietly rendering the clip on `v4` in a voice you never enrolled.
Re-enrol the voice in the [Studio](https://vieneu.io/#/clone);
the new `clone_…` id works on `POST /v1/tts` and `POST /v1/tts/stream` exactly
as before.
**What is not affected.** Audio already generated on `v3` is untouched —
library entries and download links keep working. And this is a cloud-API change
only: the on-device **v3 Turbo** in the Python SDK and the Windows app runs on
your own machine and is unchanged.
---
## 2026-09-23 — Cloning moved to the web Studio
### ⚠️ Behaviour change
**The clone-enrolment endpoints are closed.** `POST /v1/voices`,
`DELETE /v1/voices/{voiceId}`, `POST /v1/clone`, `POST /v1/upload` and
`POST /v1/prepare` now answer `410` with `code: "CLONE_WEB_ONLY"` and a
`cloneUrl`. Nothing is billed and no GPU work starts — the refusal happens
before either. Create the voice at
[vieneu.io/#/clone](https://vieneu.io/#/clone) instead.
Why the web: enrolling well is a guided job. The Studio denoises the clip,
transcribes it with Whisper and lets you hear the voice before it is saved. A
reference that is noisy, mistranscribed or the wrong length degrades **every**
later generation with that voice, not just the call that created it — and an API
caller had no way to see that before spending the enrolment fee.
**Nothing changes for generation.** A voice cloned on the web is listed by
`GET /v1/voices` when you send your key (`"kind": "cloned"`) and its `clone_…` id
remains an ordinary `voiceId` on `POST /v1/tts` and `POST /v1/tts/stream`, at the
same price as a catalogue voice. Integrations that only *use* clones need no
change at all.
The clone routes are gone from [`/openapi.json`](pathname:///openapi.json) and
the [API reference](/api-reference), so generated clients will drop them on their
next regeneration. The old clone error codes (`CLONE_MONTHLY_CAP`,
`CLONE_QUOTA_EXCEEDED`, `CLONE_REF_*`, …) belonged to those routes and can no
longer reach an API key; the same plan limits still apply in the Studio. See
[Errors → Cloning](./errors.md#cloning).
---
## 2026-09-16 — Safe retries and batch polling on `POST /v1/tts`
### New
**`Idempotency-Key` on `POST /v1/tts`.** The server creates the job and charges
the tokens *before* it answers, so a client whose request timed out — or was cut
by a `502` from the edge during a deploy — could never tell "never happened" from
"happened, answer lost". Send a client-chosen key (8–128 printable ASCII
characters; a UUID is ideal) and a retry with the same key and the same body
returns the **first** job: same `jobId`, no second charge, and the response
header `Idempotent-Replayed: true`. Keys are scoped to your API key and kept for
24 hours. Same key with a different body is refused with `422
IDEMPOTENCY_KEY_REUSED`; the same key while the first call is still running is
`409 IDEMPOTENCY_IN_FLIGHT` with `Retry-After: 1` — retry with the **same** key,
never a new one. Requests without the header behave exactly as before. See
[Retrying safely](./overview.md#retrying-safely).
**`GET /v1/tts?ids=a,b,c`** — poll up to 50 jobs with one request. Each entry
has exactly the shape of `GET /v1/tts/{jobId}`; ids that are unknown or not yours
are listed in `missing` rather than failing the batch. Not throttled, like the
single poll. Prefer it whenever more than one job is in flight: the edge proxy
caps concurrent connections per source address, and one batch call costs the
server one database read where eight single polls cost sixteen.
**`retryAfterSeconds` in error bodies, mirrored as `Retry-After`.**
`CONCURRENT_CONFLICT` now says how long to wait (one second) instead of leaving
you to guess; `QUEUE_FULL` and `USER_QUEUE_FULL` already carried it, and it is
now documented on the error body. `USER_QUEUE_FULL` — the per-account ceiling on
queued jobs — is also new to this document, though not to the API.
### Behaviour change — additive
**A full queue no longer touches your balance.** `POST /v1/tts` used to charge,
discover the queue was full, and refund inline. It now asks the queue first, so
a `QUEUE_FULL` / `USER_QUEUE_FULL` answer moves no money — which is also what its
message has always promised.
---
## 2026-09-13 — `GET /v1/usage`
### New
**`GET /v1/usage`** — your own usage, from the same per-call ledger the platform
keeps: calls by outcome, tokens, characters, seconds of audio, generation time
(average, p95), broken down per day, per route, per voice and per error code.
`scope=key` (default) is the calling key; `scope=account` merges every key on
the account. Window up to 92 days. See
[Seeing what you were billed for](./overview.md#seeing-what-you-were-billed-for).
Nothing else changed for callers. Behind it, every `/v1` call is now recorded
with its key, route, voice, cost and outcome — the reason a usage question can
be answered exactly rather than estimated.
---
## 2026-09 — Error bodies, quota codes, and the clone engine gate
### ⚠️ Behaviour change — additive, nothing removed
Read this line first: **no field was removed, renamed, or changed in type or
meaning.** All three changes below only put more into a body that was already
being sent. The only integration that can notice is one that rejects unknown
fields — a strict schema, a Go struct decoded with unknown-field checking — and
that is the reason this is filed as a behaviour change rather than under *New*.
**Every native `/v1` error body now carries `statusCode`.** It used to appear
only on the bodies that had nothing else in them: the moment a refusal gained a
machine-readable `code`, it lost `statusCode` in the same move, so gaining a
reason cost you a field you had always been able to read. One shape now, on
every error the platform envelope covers. The HTTP status line remains the
authoritative copy. `POST /v1/audio/speech` is unaffected — it answers in
OpenAI's envelope, which has no `statusCode`.
**Quota refusals now carry `code`, plus `resetAt`, `remaining` and `required`
where the deduction supplies them.** Eight refusal sites across the seven billed
routes — `/v1/tts`, `/v1/tts/stream`, `/v1/dialogue`, `/v1/dub` (two of them: the
affordability gate before any GPU work, and the charge after it), `/v1/clone`,
`/v1/srt` and `/v1/vapi/speech` — used to throw a bare sentence, so the only way
to tell "top up" from "wait an hour" was to match English prose. They
now carry `GRANT_EXPIRED`, `GRANT_TOKENS_EXHAUSTED`, `GRANT_DAILY_LIMIT`,
`GRANT_WEEKLY_LIMIT` or `CONCURRENT_CONFLICT`, and the two limits that lift on
their own carry an ISO 8601 `resetAt` saying when. `message` is unchanged, so a
client reading only that keeps working. See [Errors](./errors.md#quota-and-grant)
for what each code means and what to do about it.
**`POST /v1/voices` now gates the engine that actually runs.** Enrolment has
been `v4`-only since 2026-08-28, but the entry check was still asserting things
about `v3` — so an enrolment that named no engine could be refused because of
the state of an engine it was never going to use. The check now runs against the
engine the request will really enrol on, and only when you named one explicitly.
- *What you will see:* an enrolment that omits `engine` — the normal path —
no longer fails for reasons that have nothing to do with it.
- *What has not changed:* an explicit `"engine": "v3"` is still refused with
`400 CLONE_ENGINE_NOT_ALLOWED`, and an explicit engine your plan does not
cover is still `403`.
---
## 2026-09 — Webhooks
### New
**Webhooks for asynchronous work.** Register an `https` endpoint with
`POST /v1/webhooks`, and instead of polling `GET /v1/tts/{jobId}` you receive
`tts.job.completed` or `tts.job.failed` the moment the job ends. The event
carries everything you need to act on it without a second call, plus a download
URL that lives for one hour.
- Signed with HMAC-SHA256 in a `Vieneu-Signature: t=…,v1=…` header — the same
shape Stripe uses, so the signature verifier you have already written carries
over almost untouched. A copy-paste-and-run `verifySignature` is in
[Webhooks](./webhooks.md).
- **At-least-once and unordered delivery.** Duplicates are normal: dedupe on
`id` (stable per job and event type). Do not dedupe on the signature — every
attempt is signed afresh, so `t` and `v1` both change.
- A failed delivery is retried up to 6 attempts, at `t+0`, `+0.5m`, `+2.0m`,
`+5.5m`, `+13.0m` and `+28.5m` — about **28.5 minutes** for the whole run, and
the longest single gap between two attempts is **15.5 minutes**. (This entry
originally said "about 15 minutes"; that was arithmetic on the wrong backoff
formula, see [Webhooks](./webhooks.md#retries-and-backoff).) Size any
reconciliation threshold **above 15.5 minutes**. Once the attempts run out the
event is **dead-lettered**: it will not be sent again, a notification appears
in your account, and `GET /v1/webhooks/{id}/deliveries` tells you why. The
audio is **not** lost — `GET /v1/tts/{jobId}` still returns the job and a
fresh download URL.
- The endpoint must be `https` and must point at a public address. Private
ranges, loopback, link-local, CGNAT and reserved ranges are all refused — both
at registration and again at the moment of delivery. Any `3xx` response counts
as a failed delivery.
- Rotate the signing secret with `POST /v1/webhooks/{id}/rotate-secret`. The old
secret keeps working for 24 hours and each event carries one `v1=` per live
secret, so you can roll out the new one without dropping a single event.
Nothing changes behaviour here: polling works exactly as before, and an account
that has registered no endpoint sees no difference at all.
---
## 2026-08 — Output formats, OpenAI compatibility, and Vapi
### New
**Cloned voices work on `POST /v1/tts/stream`.** Pass a `clone_…` id (from
`GET /v1/voices`, created with `POST /v1/voices`) as `voiceId` and it streams,
just like a built-in voice. The streaming path used to refuse every `clone_…` id
while the web app had been streaming cloned voices for a long time — the worker
supported it all along; only the public surface blocked it.
- Billing is **unchanged**: still submitted characters × the engine multiplier.
Clones carry no surcharge of their own; the only charge is the one-off
enrolment fee on `POST /v1/voices`.
- A cloned voice is bound to the engine it enrolled on, so the "streaming is
`v4` only" rule still applies. New clones all enrol on `v4`.
- This is not a behaviour change: whatever worked yesterday still works exactly
the same.
**Multiple output formats.** `POST /v1/tts/stream` and `POST /v1/audio/speech`
now take `mp3`, `opus`, `pcm` (raw, headerless) and `ulaw` (G.711 8 kHz for
telephony), with a choice of sample rate. Before, WAV at 48 kHz was the only
option. mp3 is roughly 6× smaller than WAV.
**`POST /v1/audio/speech` is a real drop-in for OpenAI.** Point an OpenAI SDK at
`https://api.vieneu.io/api/v1` with your API key and it works, with nothing else
to change. `speed` is now applied, `model` selects the engine, and
`stream_format` lets you take audio while it is still being generated (chunked,
or OpenAI-style SSE). `GET /v1/audio/voices` was added so OpenAI clients can
list voices.
**`POST /v1/vapi/speech`** — the custom-voice webhook for [Vapi](https://vapi.ai).
Point your assistant's `custom-voice` at this URL and your voice agent speaks
Vietnamese. Details in [Vapi integration](./vapi.md).
**Rate-limit headers on every response.** `X-RateLimit-Limit`,
`X-RateLimit-Remaining`, `X-RateLimit-Reset`, and `X-Request-Id` — quote
`X-Request-Id` when you contact support, it is the id our logs are keyed by.
**Public documentation.** The Cloud API section on docs.vieneu.io, with
reference decoders in Python and JavaScript for streaming.
### ⚠️ Behaviour change
**`aiRefine` now defaults to OFF, not ON.**
The AI refinement step (content moderation + pronunciation normalisation for
formulas, acronyms and mixed-in English) used to run by default on `/v1` and
carried a surcharge. You now have to ask for it explicitly with
`"aiRefine": true`.
- *What you will see:* cost per request **down**, latency **down**, and your text
read **exactly as you sent it**.
- *If you want the old behaviour:* add `"aiRefine": true` to the request. On
`POST /v1/srt` (multipart) it is an `aiRefine=true` field.
- Deterministic text preparation (chemistry readings, UPPERCASE folding, the
sea-g2p phoneme layer) runs in every case — Vietnamese is still pronounced
correctly.
**The `response_format` default on `/v1/audio/speech` moves from `wav` to
`mp3`.** That is OpenAI's default, and an OpenAI client that leaves the field
empty is expecting mp3. If you need WAV, send `"response_format": "wav"`.
**Rate limits change unit: from per IP to per API key.** They used to count IP
addresses, which meant an office sharing one network had to split a limit
between them, while one key spread across several machines multiplied it. They
now count the key — the thing you actually bought. Alongside that, the ceiling
on the three main synthesis routes was **raised to 300 requests per minute** so
that nobody ends up with less than before.
This is the **application-tier** limit. In front of it sits an edge proxy that
still counts source addresses, and that tier has not changed. How to tell them
apart: a `429` carrying `X-RateLimit-*` is your key hitting the application
ceiling; a bare `429`/`503` with no headers at all is the edge proxy blocking by
IP — shared with everyone else behind that address. See
[Rate limits](./overview.md#rate-limits-and-request-ids).
### Fixed
- Streaming through `api.vieneu.io` is no longer buffered by the proxy — audio
arrives as it is generated, the way it was designed to.
- `POST /v1/tts/stream` now reports the real format and sample rate in the
`X-Output-Format` / `X-Sample-Rate` headers, instead of echoing back what you
asked for.
- Truncated streams are refunded more reliably: a stream that ends correctly but
carries **no audio at all** is now refunded too.
### Known limitations
`stream_format` (and streaming in general) should only be used with `pcm` and
`ulaw`. Audio frames are encoded independently of one another, so concatenated
mp3 has a gap of silence at every join, and opus becomes a chain of Ogg streams
that most browsers only play the first of. If you need a complete mp3 or opus
file, send an ordinary, non-streaming request.
---
# Integrations overview
Source: https://docs.vieneu.io/docs/integrations/overview
:::info The n8n node is not published yet
The n8n node (`n8n-nodes-vieneu` 0.1.0) is marked `private` and has never been
published to npm. It is built and working, but the repository it lives in is not
public, so there is no install command you can run today — **contact VieNeu to
get the package**, and watch the [changelog](../cloud-api/changelog) for the
release.
Nothing else here depends on it. The MCP server, the OpenAI-compatible endpoint,
the raw API, webhooks and Vapi are live and need nothing installed.
:::
This page routes you to the right one. Each destination owns its own detail;
this one deliberately repeats none of it.
## Which path is yours
| What you already have | How VieNeu plugs in | Where |
|---|---|---|
| An app that speaks the **OpenAI TTS protocol** — Open WebUI, SillyTavern, LobeChat, LiteLLM, anything on the OpenAI SDK | Three settings: base URL, key, voice id. No plugin, no adapter. | [OpenAI-compatible apps](./openai-clients.md) — per-client setup and troubleshooting by symptom · [endpoint reference](../cloud-api/openai-compatible) |
| An **AI assistant** — Claude, ChatGPT, Cursor, Claude Code | A hosted MCP server: add `https://api.vieneu.io/mcp`, sign in with your VieNeu account, ask for speech in plain language. Nothing to install. | [MCP server](./mcp/index.md) |
| **n8n** | A community node (self-hosted n8n, built from source), or two ready-made workflows built from n8n's own HTTP Request node that install nothing and run on n8n Cloud. | [n8n node](./n8n.md) |
| A **voice agent or phone system** | Vapi has a dedicated webhook route. LiveKit Agents, Pipecat and anything else with an OpenAI TTS plugin use `/v1/audio/speech` with `response_format: "pcm"`. | [Vapi](../cloud-api/vapi) · [LiveKit / Pipecat](./openai-clients.md#livekit-agents) · [Streaming](../cloud-api/streaming) |
| **Your own code** | Call `/api/v1` directly — that surface is much larger than the OpenAI-shaped one. | [Cloud API overview](../cloud-api/overview) · [API reference](/api-reference) |
| Something that must **react when a job finishes** | Register a webhook instead of polling. | [Webhooks](../cloud-api/webhooks) |
## What every path shares
Base URL `https://api.vieneu.io/api/v1`. One key, under either header name:
```
Authorization: Bearer vn_sk_...
X-API-Key: vn_sk_...
```
`vn_sk_` keys are live; `vn_test_` keys behave identically but cap each request
at 100 words. Any other string — `sk-…`, `none`, an empty placeholder — is a
**401 before any lookup**, so a client that insists on a non-empty key field must
be given a real one. A `Bearer` token containing a dot is read as a JWT and
ignored, which is why pasting one produces "API key required" rather than
"invalid key". A third header, `X-VAPI-SECRET`, is accepted on every `/v1` route
— last in precedence, so an explicit header always wins. It exists because a
Vapi assistant's configuration can set neither of the other two. See
[Authentication](../cloud-api/overview#authentication).
Synthesis is billed per **submitted character**, minimum 50, times the engine's
multiplier; a request that produces no audio is refunded automatically. See
[Billing](../cloud-api/overview#billing).
## One request end to end
```bash
curl -sS https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{"input": "Xin chào, đây là VieNeu.", "response_format": "mp3"}' \
--output speech.mp3
```
Name `response_format` even though `mp3` is the default. Only an *explicit* mp3
is guaranteed to be mp3: when the request never mentions a format and the worker
cannot encode one, the endpoint falls back to wav rather than failing, and
`-sS --output` discards the `X-Output-Format` header that would have told you.
Written to `speech.mp3`, those are bytes your player refuses.
Omitting `voice` uses the first active voice on the configured default engine
(see [Engines](../cloud-api/overview#engines) — the engine also sets the price
multiplier). To choose one, ask which engine is the default, then list that
engine's voices:
```bash
curl -sS https://api.vieneu.io/api/v1/engines # no key needed; find "isDefault": true
curl -sS "https://api.vieneu.io/api/v1/audio/voices?engine=" \
-H "Authorization: Bearer $VIENEU_API_KEY"
```
`v4` is today's default — and, since `v3` was retired from the cloud API on
2026-09-24, the only engine — but the registry is the source of truth, so
substitute whatever the first call reports. The retired catalogue's ids (opaque
`vieneu-…` slugs) shared almost no space with `v4`'s display names, so a voice
saved from an unfiltered list before the retirement 400s now; pick a fresh one.
Details in [Getting voice ids](./openai-clients.md#getting-voice-ids).
## What the OpenAI shape cannot reach
`POST /api/v1/audio/speech`, plus the `GET /api/v1/audio/voices` companion its
voice picker reads, is the whole compatibility surface — there is no adapter to
install and nothing to maintain per client. Two things live outside that shape,
and if you are comparing providers they are worth testing directly rather than
inferring:
- **Incremental streaming.** `POST /v1/tts/stream`, and `stream_format` on the
OpenAI route, deliver audio as it is generated rather than as one buffered
download wearing a streaming label. It is declared per engine — `features` on
`GET /v1/engines` includes `stream` for `v4`; the `v3` engine retired on
2026-09-24 never had it, because it handed over its chunks in a burst at the
end and could not honour the latency promise. See
[Streaming](../cloud-api/streaming).
- **Cloned voices, usable from code.** You clone a voice once in the web Studio
([vieneu.io/#/clone](https://vieneu.io/#/clone)) — the guided path that
denoises the clip, transcribes it and plays it back before saving — and its
`clone_…` id then works on `POST /v1/tts` and `POST /v1/tts/stream` like any
catalogue voice. It is the one thing an OpenAI-compatible client cannot reach:
a `clone_…` id on `/v1/audio/speech` is a 400.
The smaller frictions of the OpenAI shape — no `GET /v1/models`, unknown body
fields rejected rather than ignored, OpenAI voice names unmapped, the doubled
`/api/v1` base-URL trap, streaming needing two fields, browser-side CORS, and
which 429 came from where — all have a symptom and a fix in
[OpenAI-compatible apps](./openai-clients.md).
## Latency, honestly
**On the standard plans**, first audio lands in roughly **1–2 seconds**. That is
comfortable for read-aloud, IVR prompts, dubbing and assistants that tolerate a
beat before speaking. It is **not** the 200–300 ms class that hard real-time
conversational agents expect. Enterprise runs on a dedicated node where that
ceiling is agreed per contract instead — down to ≤ 250 ms. Either way, design
around the number you measure on the plan you are on — see
[Latency, honestly](../cloud-api/streaming#latency-honestly).
---
# Drop-in for OpenAI-compatible apps
Source: https://docs.vieneu.io/docs/integrations/openai-clients
A lot of software already speaks `POST /v1/audio/speech` and lets you point it at
a custom base URL. VieNeu implements that endpoint, so those apps can use
Vietnamese voices with no plugin and no adapter.
This page is the setup sheet: what to put in which field, for each client.
The VieNeu side of every section below is exact. The client side is written as
"which value goes in which kind of field", because setting labels move between
versions — **verify the field names against your build**.
## The three values
Every client needs the same three things.
| Value | What to enter |
|---|---|
| Base URL | `https://api.vieneu.io/api/v1` (see the trap below) |
| API key | your VieNeu key — starts `vn_sk_` (live) or `vn_test_` (test) |
| Model | `tts-1`, `tts-1-hd`, `gpt-4o-mini-tts`, `vieneu` or `vieneu-v4` → `v4`, the only engine since `v3` was retired on 2026-09-24. Any other name is accepted and ignored — **except** a `vieneu-…` name that is not a live engine, which is a 400. `vieneu-v3` is now one of those. |
The one carve-out in that last row is the one a VieNeu user is most likely to
trip. `vieneu-v5` or `vieneu-turbo` does **not** fall through to the default
engine the way `whatever-tts` does; it returns 400 `invalid_request_error` with
`param: "model"`, listing the engine names that do exist. The prefix is read as
an explicit engine choice, and quietly rendering it on some other engine would
bill at a rate you never picked. If you send both `engine` and `model`, `engine`
wins.
`vieneu-v3` is the case you are most likely to actually hit. It was a valid
choice until `v3` was retired from the cloud API on 2026-09-24, and now answers
400 with `Model 'vieneu-v3' is retired. Engine "v3" was retired on 2026-09-24.
Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues
share no ids.` Change the model to `vieneu-v4` (or `tts-1`) and re-check the
voice in the same pass — a `vieneu-…` slug saved from the old `v3` catalogue
will 400 on `param: "voice"` next.
A fourth field, `voice`, is optional but usually wanted: omitting it uses the
first active voice on the resolved engine. What you cannot do is reuse OpenAI's
names — `alloy` / `nova` / etc. are **not** mapped and return 400. Get real ids
from `GET /api/v1/audio/voices` (below).
### The base-URL trap
The path is `/api/v1/audio/speech`. The doubled-looking `/api/v1` is real — the
server sets a global `api` prefix *and* mounts the public API at `v1`. So the
value you type depends on what your client appends to it:
| If the field means… | Enter |
|---|---|
| "the OpenAI base — I append `/audio/speech`" | `https://api.vieneu.io/api/v1` |
| "the host — I append `/v1/audio/speech`" | `https://api.vieneu.io/api` |
| "the full endpoint URL" | `https://api.vieneu.io/api/v1/audio/speech` |
Most clients mean the first. A wrong pick is always a **404, and always JSON** —
but in VieNeu's platform error shape rather than the OpenAI envelope this
endpoint otherwise uses:
```json
{ "statusCode": 404, "message": "Cannot POST /api/v1/v1/audio/speech", "error": "Not Found" }
```
Two things to read off it:
- The `Cannot POST …` message and the **absence** of the
`{"error":{"message","type","param","code"}}` wrapper that every real
`/v1/audio/speech` failure carries mean the URL shape is wrong, not the key.
Do not go hunting your API key on this one.
- The path inside the message is the URL your client actually built. Compare it
to `/api/v1/audio/speech` and the difference tells you which row above you
needed — the example here doubled `/v1`, so that client wanted
`https://api.vieneu.io/api`.
### There is no `/v1/models`
VieNeu serves TTS only. There is no `/v1/models` route and no
`/v1/chat/completions`. Two consequences:
- A **"Test connection" / "Verify" button that probes `/models` will report
failure** even though synthesis works. Ignore it and send a real request.
- A **model dropdown fed from `/models` will be empty.** Type the model name by
hand.
In Open WebUI, SillyTavern and LobeChat, this base URL belongs **only** in the
audio/TTS provider setting — never in the general OpenAI/LLM endpoint setting, or
the chat model breaks.
### Prove it works first
Before touching any client, confirm the base URL and the key with curl. Send
`$VIENEU_API_KEY` — your own `vn_sk_…` or `vn_test_…` key — and **no `voice`**,
so nothing but those two values can fail:
```bash
curl https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu." }' \
--output speech.mp3
```
Omitting `voice` makes the server pick the first active voice on the resolved
engine, which is what keeps this step honest: a voice id you have not verified
yet would 400 on `param: "voice"` and leave you unable to tell a bad voice from
a bad key or a bad base URL. Pick a real voice in the next step, once this one
has passed.
If that produces playable audio, everything after this is client configuration.
### Getting voice ids
```bash
# Which engine will my requests land on? No API key needed.
curl https://api.vieneu.io/api/v1/engines
# Voice ids for that engine. API key IS required here.
curl "https://api.vieneu.io/api/v1/audio/voices?engine=v4" \
-H "Authorization: Bearer $VIENEU_API_KEY"
```
`GET /api/v1/engines` returns one entry per enabled engine with `key`,
`isDefault`, `billingMultiplier` and `features`. Use the one with
`isDefault: true` — since `v3` was retired on 2026-09-24 that is the only entry,
`v4`.
:::warning Voice ids saved before 2026-09-24 may be dead
`audio/voices` now lists `v4` only, filtered or not, and `?engine=v3` is a 400.
But the retired V3 catalogue's ids were `vieneu-…` slugs, V4 ids are display
names (`Ngọc Lan`), and the two shared almost no id space — so a voice a client
saved from an unfiltered list before the retirement will 400 on
`param: "voice"` now. Refresh the picker from the call above, and keep passing
`?engine=v4`: it costs nothing and keeps the request explicit.
:::
Voice matching is case-insensitive but **not** diacritic-insensitive, and V4 ids
contain spaces. A client that slugifies or strips accents before sending will
400 on every request.
Now re-run the smoke test with an id from that list, copied exactly as returned:
```bash
curl https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu.", "voice": "PASTE_AN_ID_HERE" }' \
--output speech.mp3
```
If the first curl worked and this one 400s on `param: "voice"`, the id is the
only thing that changed — check the engine filter and the accents before
anything else.
:::note Cloned voices do not work here
`clone_…` ids are rejected on `/v1/audio/speech` with a message naming the two
routes that accept them: `POST /v1/tts` and `POST /v1/tts/stream`. No
OpenAI-compatible client can reach a cloned voice.
:::
## Open WebUI
Audio/TTS settings, OpenAI engine:
| Field | Value |
|---|---|
| TTS engine / provider | the OpenAI-compatible option |
| API base URL | `https://api.vieneu.io/api/v1` |
| API key | `vn_sk_…` |
| Model | `tts-1` |
| Voice | a real id from `audio/voices`, e.g. `Ngọc Lan` |
Notes:
- Put this in the **audio** settings, not the model/connection settings.
- The voice field must accept free text. If your build only offers OpenAI's six
names, they will all 400 — verify on your build.
- Response format: leave it at the default. Open WebUI plays mp3, which is also
VieNeu's default.
## SillyTavern
TTS extension, OpenAI-compatible provider:
| Field | Value |
|---|---|
| Provider | the OpenAI TTS option |
| API / base URL | `https://api.vieneu.io/api/v1` |
| API key | `vn_sk_…` |
| Model | `tts-1` |
| Voice (per character) | a real id from `audio/voices` |
Notes:
- SillyTavern assigns a voice per character. Every one of them must be a VieNeu
id — a character left on `alloy` fails while the others work, which reads as a
flaky integration.
- If the extension only offers a fixed voice dropdown rather than a text field,
it cannot address VieNeu voices. Verify on your build.
- If your SillyTavern runs its TTS call from the **browser** rather than its own
server, see [Browser-side clients](#browser-side-clients).
## LobeChat
TTS / audio settings, OpenAI provider:
| Field | Value |
|---|---|
| OpenAI TTS base URL / proxy URL | `https://api.vieneu.io/api/v1` |
| API key | `vn_sk_…` |
| Model | `tts-1` |
| Voice | a real id from `audio/voices` |
Notes:
- LobeChat keeps separate settings for the chat provider and the TTS provider.
This URL goes in the **TTS** one only.
- LobeChat is commonly deployed so that the browser calls the TTS provider
directly — see [Browser-side clients](#browser-side-clients) before you debug
anything else.
## LiteLLM
Add VieNeu as a model in `config.yaml`:
```yaml
model_list:
- model_name: vieneu-tts
litellm_params:
model: openai/tts-1
api_base: https://api.vieneu.io/api/v1
api_key: os.environ/VIENEU_API_KEY
```
Then, with the proxy running, call it exactly as you would OpenAI:
```bash
curl http://localhost:4000/v1/audio/speech \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "vieneu-tts", "input": "Xin chào", "voice": "Ngọc Lan" }' \
--output speech.mp3
```
Notes:
- **`LITELLM_MASTER_KEY` is not your VieNeu key.** It is LiteLLM's own proxy
credential, whatever you set when you started the proxy — the port
(`:4000` above) is LiteLLM's too. Your `vn_sk_…` key appears only in
`config.yaml`, reached through `os.environ/VIENEU_API_KEY`; the proxy is what
attaches it to the upstream call.
- The `openai/` prefix on `model` tells LiteLLM to pass the request through in
OpenAI's shape, which is what VieNeu answers. The name after it (`tts-1`) is
what reaches VieNeu.
- The VieNeu values above are exact; the surrounding LiteLLM keys are its
standard `model_list` shape — confirm them against your LiteLLM version.
- **Do not enable any request-enrichment that adds body fields.** VieNeu rejects
unknown fields with 400 rather than ignoring them — see
[400](#400-invalid_request_error).
## LiveKit Agents
The OpenAI plugin takes a base URL and a key:
```python
from livekit.plugins import openai
tts = openai.TTS(
model="vieneu-v4",
voice="Ngọc Lan",
base_url="https://api.vieneu.io/api/v1",
api_key="vn_sk_...",
)
```
Notes:
- **Pin `vieneu-v4`, or omit `model`.** Either lands on `v4`, the only engine
since `v3` was retired on 2026-09-24. `vieneu-v3` — which this endpoint used
to accept with streaming, answering 200 and delivering v3's chunks in a burst
at the end with no error telling you so — is now a 400 with `param: "model"`
saying the engine was retired.
- `GET /v1/engines` returns the live billing multiplier; see
[Engines](../cloud-api/overview#engines).
- For streaming, VieNeu needs **both** `response_format: "pcm"` and
`stream_format`. `response_format` defaults to mp3, so setting only
`stream_format` is a 400. Whether the plugin sends either, and whether it lets
you add them, is the thing to verify on your build. Without them the call
still works — it just returns a complete mp3 instead of a stream.
- `pcm` is headerless and defaults to **24000 Hz** on this route (the engines'
native rate is 48000). Read the actual rate from the `X-Sample-Rate` response
header and configure the pipeline to match, or the audio plays at the wrong
pitch.
## Pipecat
```python
import os
from pipecat.services.openai.tts import OpenAITTSService
tts = OpenAITTSService(
api_key=os.environ["VIENEU_API_KEY"],
base_url="https://api.vieneu.io/api/v1",
model="vieneu-v4",
voice="Ngọc Lan",
)
```
Notes:
- The import path for Pipecat's OpenAI TTS service has moved between releases —
check yours. The constructor arguments and the values above are what matter.
- The same three streaming rules as LiveKit apply: pin `vieneu-v4`, send
`response_format: "pcm"` **and** `stream_format`, and take the rate from
`X-Sample-Rate` rather than assuming 48 kHz.
- If your version hard-codes a `response_format` VieNeu does not accept
(`aac`, `flac`), the request 400s naming the field. `wav`, `mp3`, `opus`,
`pcm` and `ulaw` are the accepted values.
## Browser-side clients
In production, VieNeu's CORS policy allows only VieNeu's own origins. A client
that calls `/v1/audio/speech` **from the page** rather than from its own server
fails at the preflight, with no response body to explain it.
Whether a given LobeChat or SillyTavern deployment does its TTS server-side is a
per-deployment question — verify on your build. If it is browser-side, put a
proxy of your own in front (LiteLLM works well for this) and point the client at
that.
Browser JavaScript can read `X-Request-Id`, `X-Sample-Rate`, `X-Output-Format`,
the `X-RateLimit-*` headers and `Retry-After`; everything else is hidden by the
browser. (`X-Stream-Format` is CORS-exposed too, but do not write a read for it
here — only the native `POST /v1/tts/stream` sends it. On this route the framing
is already unwrapped for you, so the header is always absent.)
## Troubleshooting by symptom
### No audio
Work down this list — each has a different cause.
- **The saved file contains JSON.** On the raw streaming path, a generation that
produced nothing answers **502** with an error body where audio was expected. A
client that writes the response body straight to a `.pcm` file ends up with
JSON in it. Check the first bytes of the file.
- **The file downloads instead of playing.** The synchronous response carries
`Content-Disposition: attachment`. Clients that fetch the body are unaffected;
one that navigates to the URL gets a download.
- **Audio plays at the wrong pitch or speed.** `pcm` and `ulaw` are headerless —
the bytes carry no sample rate. `pcm` is 24000 Hz here unless you asked
otherwise; `ulaw` is always 8000. Read `X-Sample-Rate` instead of assuming.
- **The player refuses an `.mp3`.** If you omitted `response_format` and the
worker serving you cannot encode mp3, VieNeu falls back to **WAV bytes** rather
than failing. Read `X-Output-Format` to see what you actually got. (Had you
asked for mp3 explicitly, you would have got a 503 naming the format instead —
an explicit choice is never silently substituted.)
### 400 `invalid_request_error`
The `param` field in the error body names the culprit for most of these — with
one exception, and it is the first cause listed below. The common causes:
- **An unknown body field.** VieNeu rejects fields it does not declare rather
than ignoring them. The accepted set is exactly: `model`, `input`, `voice`,
`response_format`, `sample_rate`, `speed`, `instructions`, `stream_format`,
`emotion`, `aiRefine`, `engine`. Note there is **no `stream` boolean** — a
client that sends `stream: true` (as for chat completions) gets a 400. So does
`user`, `language`, or any client-specific extension. This is the most common
reason a client that works against OpenAI fails against VieNeu: inspect the
request body it actually emits.
**Read `message`, not `param`, for this one.** The rejection comes from the
request validator rather than from the handler, so `param` comes back as the
literal string `"property"` and never the offending field name. The name is in
the message: `property stream should not exist`. Only the first offending
field is reported per request, so if your client adds several, fix and retry
until it passes.
- **An unrecognised `vieneu-…` model name.** `param: "model"`. Other unknown
model names are ignored; this prefix is not.
- **An OpenAI voice name.** `alloy`, `echo`, `fable`, `onyx`, `nova`, `shimmer`
are not mapped. `param: "voice"`.
- **A cloned voice id.** `clone_…` works only on `POST /v1/tts` and
`POST /v1/tts/stream`.
- **`stream_format` with a non-streamable format.** Only `pcm` and `ulaw` can be
streamed, and `response_format` defaults to **mp3** — so setting
`stream_format` alone is always a 400.
- **`aac` or `flac`.** Real OpenAI formats VieNeu cannot encode; the message says
so explicitly.
- **A `sample_rate` that contradicts the format.** `opus` is always 48000 and
`ulaw` always 8000; a conflicting rate is rejected, not overridden. Valid rates
are 8000, 16000, 22050, 24000, 44100, 48000.
- **`speed` outside 0.25–4.0.** Inside that range it is clamped to 0.5–2.0 and
never errors; outside it, validation rejects it.
- **A test key over 100 words.** `vn_test_` keys cap each request at 100
whitespace-separated words. The message names your actual word count.
### 401 `authentication_error`
- **`Invalid API key format.`** — the key must literally begin `vn_sk_` or
`vn_test_`. This is checked before any lookup, so placeholder keys some clients
ship or require (`sk-...`, `none`, `ollama`) fail here. If the client refuses
to save an empty key field, it needs a real VieNeu key.
- **`API key required.`** although you set one — a bearer token containing a dot
is treated as a JWT and ignored as an API key. Real VieNeu keys are hex and
never contain a dot, so this means something wrapped or replaced your key.
- **`Invalid or revoked API key.`** — a revoked key is indistinguishable from an
unknown one.
Keys go in `Authorization: Bearer ` or `X-API-Key: `; both work. If
your client sends both with different values, `X-API-Key` silently wins.
### 403 — usually "out of credit" {#403-out-of-credit}
**Running out of tokens is 403 here, not 402.** Nothing in the public API emits
402, so a client that watches for it never learns it has stopped being paid for.
Watch 403 instead. Three messages, all typed `insufficient_quota`:
- `Insufficient tokens. You have tokens remaining but this request costs tokens.`
— the ordinary end of a grant, and the normal first failure once a trial
allowance empties.
- `Your API token package has expired. Please renew your Developer plan.` —
time, not consumption. Buying more tokens does not fix it; renewing does.
- `API token grant is not active. Check your Developer plan status.`
Two other 403s are *not* about credit, and the `type` field tells them apart:
- `Your plan does not include engine "v4". Upgrade your Developer plan to
use it.` — typed `invalid_request_error`, `param: "engine"`. Raised only when
you pinned an engine explicitly and your plan does not cover it; with `v4`
the only engine since 2026-09-24 you should rarely see it — drop
`model`/`engine` and the default engine works.
- `This feature is not available yet.` — typed `authentication_error`, from a
feature gate rather than from billing.
So do not branch on `type` alone to detect "out of credit", and do not branch on
403 alone either. Both together are unambiguous.
A daily or weekly cap is **429**, not 403 — see below.
### 429 `rate_limit_exceeded`
Read the headers to tell the sources apart:
| Headers present | Source | Counted against |
|---|---|---|
| `Retry-After` + `X-RateLimit-*` | VieNeu's throttler on synthesis routes — 300 requests/min **by default** | your API key |
| none of them, non-JSON body | the edge proxy — 3000 requests/min | your **source IP**, pooled with everyone behind it |
| JSON, message contains `Resets at ` | your token grant's daily or weekly cap | your grant |
Back off in all three cases. The middle one is worth knowing about: if your
client runs on shared or NAT'd egress, you can trip a limit no amount of tuning
on your side will fix.
**Do not hard-code 300.** It is a deployment setting
(`PUBLIC_API_SYNTHESIS_RPM`), not a constant, and other pages quote other
figures because they were written against other deployments. The row's first
column is the durable part: read your live budget from `X-RateLimit-Limit` and
`X-RateLimit-Remaining` on any successful response.
The third row is the one people mistake for the first. A daily or weekly grant
cap is a 429 whose message ends `Resets at ` — no amount of
slowing down clears it before that time. Running out of tokens outright is
[403](#403-out-of-credit), never 402.
### 503 — the fleet, not your request
Nothing about your request needs changing; retry shortly. Four causes:
- **No worker for the engine you pinned.** `No V4 TTS worker is available right
now. Please try again shortly.` (the streaming path says `…available for
streaming right now`.) Raised only for a **non-default** engine: a request
that takes the default engine falls back to the configured worker rather than
503, so dropping `model`/`engine` is a valid workaround here.
- **No active voice on that engine**, when you omitted `voice`. `No voice is
currently available on engine "v4". Pass an explicit voiceId from GET
/v1/voices.` This one arrives as a 503 but typed `invalid_request_error` with
`param: "voice"` — a catalogue problem wearing a request-shaped label. It is
raised before billing, so nothing was charged.
- **A format you asked for explicitly that the worker cannot encode.** Had you
left `response_format` off you would have got WAV bytes instead; an explicit
choice is never silently substituted. See [No audio](#no-audio).
- **A `sample_rate` the worker cannot produce**, on the streaming path. The
message names both what you asked for and what the worker offered.
Every one of these that got as far as billing refunds automatically — you are
not charged for audio you did not receive.
### The stream stops mid-sentence
- **With SSE, wait for the terminal event.** There is no `data: [DONE]`
sentinel. Exactly one `speech.audio.done` (audio complete, carries `usage`) or
`speech.audio.error` (it is not) always arrives. If neither did, the stream was
truncated — discard the audio; the request is refunded automatically.
- **On the raw stream path there is no such signal.** A truncated stream is
indistinguishable from a short one. If you need proof of completeness, use
`stream_format: "sse"` or the native
[`POST /v1/tts/stream`](../cloud-api/streaming).
- **It "hangs", then dumps everything at once.** That used to be `vieneu-v3`
with `stream_format`: the endpoint accepted it, returned 200, and v3 emitted
its chunks in a burst at the end. Since `v3` was retired on 2026-09-24 that
request is a 400 instead, so if you still see this pattern the buffering is
happening in a proxy or client of your own, not on the engine. Pin
`vieneu-v4` and check what sits between you and the API.
- **A non-streaming call times out on long text.** `input` has **no length cap**
on this endpoint — the request simply blocks for the full synthesis time, and
the client's own HTTP timeout gives up first. The ceilings are a 2 MB request
body and a 600s proxy read timeout. For long text use the asynchronous
`POST /api/v1/tts` + `GET /api/v1/tts/{jobId}` path instead.
Every response carries `X-Request-Id`. Quote it in a support request.
## See also
- [Integrations overview](./overview.md) — which route to take if your tool is
not on this page.
- [OpenAI-compatible TTS endpoint](../cloud-api/openai-compatible) — the full
field-by-field contract for `/v1/audio/speech`: every accepted field, the
format and sample-rate rules, and the complete error list.
- [Streaming](../cloud-api/streaming) — both streaming endpoints compared,
and the terminal-event semantics summarised above.
- [Rate limits and request ids](../cloud-api/overview#rate-limits-and-request-ids)
— the headers, and why limits count per key.
- [API reference](/api-reference) — generated from the server's own OpenAPI spec.
---
# n8n node
Source: https://docs.vieneu.io/docs/integrations/n8n
:::info Release status: built, not yet published
The VieNeu community node is finished and tested, but **the package has not been
released publicly**. It is marked `private: true` in its `package.json` and has
never been pushed to npm, so there is nothing for you to install today from a
public registry.
If you want to run it now, **contact VieNeu** and ask for the package — we can
hand you a build. Watch the
[Cloud API changelog](../cloud-api/changelog) for the announcement when it
reaches npm; the [Install](#install) section below has both paths.
Everything else on this page — the operations, parameters, output fields and
error behaviour — is accurate against the built package, so you can decide
whether the node is what you want before asking for it.
:::
VieNeu ships an n8n community node — one node, four resources, audio out as a
binary field. It is a **programmatic node**, not a trigger: it sits in the middle
of a workflow and turns Vietnamese text into a file the next node can send,
store or upload.
There are two ways to call VieNeu from n8n:
- **The node** — a voice picker with search, automatic routing between the short
and long text endpoints, and API errors mapped onto n8n's machine-readable
`failure.cause`. Needs the package (see above) and a self-hosted n8n.
- **Plain HTTP Request nodes** — nothing to install, works on n8n Cloud, and
available to everyone today. See
[Don't want to install the node?](#dont-want-to-install-the-node) at the
bottom. Two importable templates ship with the package.
:::caution Do not `npm install n8n-nodes-vieneu`
The name is **unclaimed on npm**, so whatever that command resolves to today is
not this package. Do not install it and do not paste it into n8n's community-node
dialog.
:::
## Install
Self-hosted n8n only. n8n Cloud installs verified packages from npm, so the
community-node path is unavailable there whatever happens — use the
[HTTP Request route](#dont-want-to-install-the-node) instead.
### When the package is published
This is the shape the install will take. **None of it works yet** — the package
is not on npm:
```text
Settings → Community nodes → Install → n8n-nodes-vieneu
```
Nothing below will be needed then. Check the
[changelog](../cloud-api/changelog) before following the manual path.
### Today
Ask VieNeu for the package. What arrives is the source directory
`n8n-nodes-vieneu/`, which you build yourself; the repository it lives in is
private, so there is no `git clone` to give you.
Prerequisites:
| Requirement | Where it comes from |
| --- | --- |
| **Node 20.19 or newer** | `engines.node: ">=20.19"` in the package's `package.json` |
| **pnpm** — `corepack enable` is enough | the package pins `packageManager: "pnpm@10.24.0"`; npm and yarn are not substitutes here (see below) |
| **n8n bundling `n8n-workflow` 2.37.1 or newer** | the version the package is compiled against, and the source imports `NodeConnectionTypes`, which older releases do not export |
n8n does not publish a plain "minimum version" number for this; check what your
instance actually bundles with `npm ls n8n-workflow` where n8n is installed. The
package declares `n8nNodesApiVersion: 1`. On an n8n too old to export
`NodeConnectionTypes` the node does not appear in the panel at all; on one too
old for the `Failure` type, the `failure.cause` branching documented
[below](#how-failures-surface) does nothing.
Build it:
```bash
cd n8n-nodes-vieneu
pnpm install
pnpm build
pnpm verify
```
`pnpm build` is mandatory. `dist/` is git-ignored, and `dist/` is exactly what
the `n8n` block in `package.json` points n8n at — a fresh copy has none.
`pnpm verify` is the check worth running. It `require()`s the built `dist/` the
way n8n's loader does and instantiates the exported class, which catches the
three failures that otherwise show up only as a node that silently never appears
in the panel: a `package.json` path that no longer matches the emitted layout, an
icon `tsc` did not copy, and a class n8n cannot construct.
Use `pnpm`, not `npm install`. The package sits outside the repository's own
workspace and carries an empty `pnpm-workspace.yaml` to make its directory a
workspace root; without pnpm reading that file, an install run from here walks up
to the repository root and installs nothing at all, exiting 0 with nothing to
build from.
Then either copy the build in, or point n8n at the package directory:
```bash
# Option A — copy into n8n's private-node folder, then restart n8n.
mkdir -p ~/.n8n/custom
cp -r dist/nodes dist/credentials ~/.n8n/custom/
# Option B — leave it in place and point n8n at it.
export N8N_CUSTOM_EXTENSIONS="/abs/path/to/n8n-nodes-vieneu"
n8n start
```
Two mistakes that cost an afternoon:
- `~/.n8n/custom` is **not** `~/.n8n/nodes`. The latter holds npm-installed
community nodes; a private build placed there is ignored.
- `N8N_CUSTOM_EXTENSIONS` takes a **semicolon**-separated list of absolute paths
on every platform. A colon-separated list does not error — it loads nothing.
For Docker, mount the built package into the container and set that same
variable to the in-container path.
The node should then appear as **VieNeu** in the node panel. Its full type is
`n8n-nodes-vieneu.vieneu`.
:::caution This install path has not been exercised against a live n8n
The package compiles, its unit tests pass, and `pnpm verify` loads the built
output in the shape n8n's loader expects — but nobody has yet loaded it into a
running n8n instance, and the same is true of the two workflow templates, which
are validated structurally rather than by importing them. Treat your first
install as a smoke test, and tell us what breaks rather than assuming it is
something you did.
:::
:::note Not an AI Agent tool
The node deliberately does not declare `usableAsTool`, so an AI Agent cannot
call it. Speech generation spends the account's tokens on every call, and an
agent invoking it speculatively is the wrong first experience. Agents have their
own supported surface: the [MCP server](./mcp/index.md).
:::
## Credential
One credential type, **VieNeu API**, with two fields and nothing else.
| Field | Name in workflow JSON | Required | Default |
| --- | --- | --- | --- |
| API Key | `apiKey` | Yes | — (password-masked, placeholder `vn_sk_…`) |
| Base URL | `baseUrl` | Yes | `https://api.vieneu.io/api/v1` |
**The `/api/v1` prefix in the Base URL is real.** The controller is mounted at
`v1` but the application sets a global `api` prefix, so the documented `/v1/...`
paths are served at `/api/v1/...`. A base URL of `https://api.vieneu.io/v1`
404s on everything. Change this field only to point at staging or a self-hosted
deployment; trailing slashes are trimmed, and an empty value falls back to the
default.
Press **Test** after pasting a key. The test issues `GET /voices` against the
configured base URL — it costs no tokens and is exempt from the rate limiter, but
its guard still rejects a malformed or revoked key with 401, so a wrong key fails
here instead of quietly returning the public catalogue.
Which keys are valid, and what a `vn_test_` key can do, is the same everywhere:
see [Authentication](../cloud-api/overview#authentication). The one thing that
matters in n8n is that a `vn_test_` key's 100-word-per-request cap surfaces as a
**400** on long text, not as a quota error — the node does not check the prefix
or the word count locally.
### The key never reaches a workflow variable
This is the reason the key is a credential rather than a node parameter:
- Node parameters are written into execution history, into the workflow JSON
people paste into issues, and into log lines. A credential is encrypted at
rest and referenced by id — it is not part of an exported workflow.
- No code in the package ever reads `apiKey`. n8n builds
`Authorization: Bearer {{$credentials.apiKey}}` itself, inside
`httpRequestWithAuthentication`, from the credential's `authenticate` block.
There is no variable holding the key that an expression or a log line could
reach.
- Every error message, description and attached response body is run through a
redactor first. Bearer tokens and anything shaped like `vn_sk_…`, `vn_test_…`
or `sk_…` become `***REDACTED***` before n8n stores them in execution data.
## Resources and operations
| Resource | Operation | Calls | Parameters |
| --- | --- | --- | --- |
| Speech | Generate | `POST /audio/speech` or `POST /tts` + poll | Text, Engine, Voice, Put Output File in Field, Options |
| Job | Get | `GET /tts/{jobId}` | Job ID, Download Audio, Put Output File in Field |
| Voice | Get Many | `GET /voices` | Engine, Search, Return All, Limit |
| Engine | Get Many | `GET /engines` | none |
Engine: Get Many has no parameters of its own. It returns one item per engine,
including the **live** `billingMultiplier` — read that rather than any table,
including the snapshot in [Engines](../cloud-api/overview#engines).
### What the node does not do
Not everything the API offers is exposed, and the omissions are deliberate: an
operation that spends real money or needs a file upload is a bad thing for a
workflow to reach by accident, and a Loop Over Items reaches things many times.
| Missing | Why |
| --- | --- |
| **Voice cloning** (`POST /voices`, `/clone`, `/prepare`, `/upload`) | A flat 5,000-token charge before any multiplier (15,000 on v4) for one clip, a multipart upload, and a slot against the plan's clone limit. Clones you already own *are* selectable in the Voice list; you just cannot create one from a workflow. |
| **`DELETE /voices/:voiceId`** | Destructive, and irreversible from inside a workflow. |
| **Dubbing and SRT** (`/dub`, `/srt`) | Both upload-shaped, both billed per call. |
| **`POST /dialogue`** | Key-only and no upload, but 50 turns × 5,000 characters is uncapped in the request schema — by a wide margin the largest single-call exposure on the surface. |
| **Streaming** (`POST /tts/stream`) | Its response is a custom length-prefixed frame format with heartbeat frames mixed in; n8n has no frame parser and would hand back one opaque buffer. It also caps at 4 concurrent streams per key, and every stream bills at the v4 multiplier whatever engine was sent. See [Streaming](../cloud-api/streaming) if you need that surface. |
### Picking a voice
The Voice field is a resource locator with two modes:
- **From List** — searchable, backed by `GET /voices`. Filtering happens in the
node, because the API has no search parameter: it folds diacritics (`đ`/`Đ`
included) and requires every word to match against id, name, description,
gender, region, engine and kind. Rows read
`Name — gender · region · engine (id)`, with `· your clone` inside the facet
group for **any** voice the API returns as `kind: "cloned"` — your own clones
and admin-published ones alike, so a public clone somebody else created is
labelled that way too. The list pages 250 at a time.
- **By ID** — for expressions. Copy the id exactly, diacritics included.
Setting **Engine** first scopes the list, and on Speech: Generate the list
re-loads when Engine changes. Do that: a voice renders only on its own engine and
is never substituted — a mismatch is a hard 400. The dropdown is loaded live from
`GET /engines`, so since `v3` was retired on 2026-09-24 it offers `v4` alone; the
only mismatch left is an old `v3` slug pasted By ID.
:::warning A `clone_…` voice does not work on short text
The short-text route the node picks by default, `POST /audio/speech`, rejects
every cloned voice outright — its voice check runs without a user id, so it
cannot tell your clone from anyone's and refuses them all with a 400 reading
*"Cloned voice … can only be used on POST /v1/tts or POST /v1/tts/stream."* This
applies to admin-published clones too.
The long-text route is `POST /tts`, which does accept clones — your own and
published ones. So to use a `clone_…` voice from the node, keep the item **over**
the Synchronous Route Threshold, or set **Options → Synchronous Route Threshold**
to `1` to force every item onto the async route.
Do not follow the node's 400 advice here. It says the usual cause is a voice
belonging to the other engine, which is true of catalogue voices and has nothing
to do with this.
:::
Ids are not descriptive. Many are opaque slugs such as `vieneu-2-000494` whose
voice is named `Minh Đức`, and a few are an unrelated person's name outright.
Never infer gender or identity from an id.
## Speech: Generate
| Parameter | Type | Default | Notes |
| --- | --- | --- | --- |
| Text | string (4 rows) | — | Required. Billed per character with a 50-character minimum — see [Billing](../cloud-api/overview#billing). |
| Engine Name or ID | options | account default | Empty = the account default decides, which also decides the rate. |
| Voice | resource locator | engine default | Scoped by Engine. |
| Put Output File in Field | string | `data` | Required. The binary property the audio lands in. |
| Options | collection | `{}` | Nine options, below. |
### Options
| Option | Name | Default | Notes |
| --- | --- | --- | --- |
| AI Refine | `aiRefine` | `false` | Runs AI normalisation and moderation first. Costs a surcharge on top of the character count and adds latency. No `/v1` route reports the surcharge — the node's help text cites 1.3× at the time of writing. It is also the only thing that can produce a 422. |
| Emotion Name or ID | `emotion` | engine default | Loaded from `GET /emotion-tags` for the selected engine. Engine-specific: v4 has no styles, so the list stays empty there. |
| File Name | `fileName` | derived | Empty derives a name from the format the API actually produced. |
| Format | `format` | `wav` | `wav`, `mp3`, `opus`, `pcm`, `ulaw`. |
| Download Audio | `downloadAudio` | `true` | Long-text route only — see below. |
| Job Timeout (Seconds) | `jobTimeout` | `600` | Minimum 5. How long to poll a long-text job. |
| Poll Interval (Seconds) | `pollInterval` | `2` | Minimum 0.5, and clamped to 500 ms in code regardless of what the workflow JSON says. |
| Speed | `speed` | `1` | 0.5–2.0. |
| Synchronous Route Threshold | `syncMaxChars` | `500` | Minimum 1. Where the route split happens. |
The node sends no sample rate and exposes no option for one.
:::tip You may not need Job Timeout at all
Job Timeout exists because the node polls. If your jobs are long enough that the
timeout is a real risk, register a webhook instead: VieNeu POSTs a signed event
when the job reaches a terminal state, and an n8n **Webhook** trigger node
receives it. `POST /tts` — the route this option governs — is the only route that
produces those events. See [Webhooks](../cloud-api/webhooks).
:::
### Two routes, chosen by text length
| | Short text | Long text |
| --- | --- | --- |
| Condition | length ≤ `syncMaxChars` (default 500) | longer |
| Endpoint | `POST /audio/speech` (blocking) | `POST /tts`, then poll `GET /tts/{jobId}` |
| Body fields | `input`, `response_format`, `speed`, and `voice` / `engine` / `emotion` / `aiRefine` when set | `text`, `speed`, and `voiceId` / `engine` / `emotion` / `aiRefine` when set |
| Formats | all five | WAV only — the route has no format parameter |
| Cloned voices | rejected with 400 | accepted (yours and admin-published) |
| `Download Audio` | ignored; bytes come back inline | applies |
| Extra JSON out | `sampleRate`, `requestId` | `jobId`, `duration`, `audioUrl`, `audioUrlExpiresIn` |
`format` and `mimeType` are set on **both** routes and are listed with the shared
fields [below](#output) — do not treat their presence as a signal of which route
an item took. Read `route` for that.
Note the field names differ between the routes (`input`/`voice` versus
`text`/`voiceId`). The node handles that; it matters only if you copy a body
between the node and an HTTP Request node.
Raising `syncMaxChars` makes the node hold an HTTP connection — and an n8n worker
slot — open for the whole synthesis.
The finished audio on the long-text route is fetched from a presigned S3 URL
**without** authentication. S3 rejects a presigned request that also carries an
`Authorization` header, so that one call deliberately skips the credential.
### Checks that run before anything is billed
These fail locally, with no request sent and no money spent:
- empty text;
- text over 50,000 characters;
- speed outside 0.5–2.0;
- a non-WAV format combined with text over the threshold. The long-text route
can only return WAV, so this is refused rather than silently downgraded.
### Output
Every item's JSON carries:
| Field | Meaning |
| --- | --- |
| `route` | `sync` or `async` — the only reliable way to tell which route ran |
| `engine`, `voiceId` | what was requested, or `null` for the default |
| `textLength` | raw character count |
| `billedCharacterBasis` | NFC character count with a 50-character floor |
| `billingNote` | states that the basis is before the engine multiplier and the AI-refine surcharge |
| `format` | the format actually produced; always `wav` on the long-text route |
| `mimeType` | read off the response, not the request — see below |
`billedCharacterBasis` is the **basis** of the charge, not a token total. The
backend applies four factors in this order: the character basis (50-character
floor), then the AI-refine surcharge if it is on, then the engine multiplier,
then a **global token multiplier** — an administrator setting, 1 by default, that
no `/v1` route reports. That last factor is why no number the node prints is a
quote. Read the engine multiplier live from Engine: Get Many, and treat any
budget you compute as approximate.
## Where the audio comes out
The audio is attached to the binary property named by **Put Output File in
Field** — `data` unless you change it. The same item's JSON also gains `fileName`
and `fileSize`.
Rename the field when a node upstream already occupies `data`, or when a
downstream node expects a specific name.
The mime type and file name are read off the **response**, never from the
request:
1. the `X-Output-Format` header, if it names a known format;
2. otherwise the `Content-Type`;
3. the file name comes from `Content-Disposition` when present, else
`speech.` — `wav`, `mp3`, `ogg` for Opus, and `raw` for the headerless
`pcm` and `ulaw`.
That indirection is deliberate: the short-text route can fall back to WAV against
an older worker, and these headers are the only signal that it did. Labelling the
binary from the request would hand a downstream node WAV bytes marked `audio/mpeg`.
On the long-text route the default name is `speech-.wav`.
**Feeding the next node.** Any node that consumes a file asks for the binary
property by name: give it `data`, or whatever you renamed the field to. The size
and name are also on the item's JSON as `{{ $json.fileSize }}` and
`{{ $json.fileName }}`, so a downstream IF can check them without touching the
bytes. To keep a large WAV out of workflow memory when you only need the link,
switch **Download Audio** off on the long-text route and pass
`{{ $json.audioUrl }}` along instead — it stays valid for a limited time, and the
JSON reports how long as `audioUrlExpiresIn`.
## How failures surface
Every API failure becomes a `NodeApiError` whose message reads:
```text
VieNeu API error 403 while synthesizing speech — (request id )
```
The prose is in `description`. The part to branch on is n8n's machine-readable
`failure.cause`. What each status means on the API side is in
[Errors](../cloud-api/overview#errors); this table is the mapping the node
adds on top:
| Status | `failure.cause` | What happened |
| --- | --- | --- |
| 400 | `configuration-invalid` | Usually a voice belonging to the other engine — but also a `clone_…` voice on the short-text route, non-Vietnamese text, and a `vn_test_` key over its 100-word cap. |
| 401 | `credential-invalid` | The key was rejected. Paste a current one into the credential. |
| 403, message mentions an engine | `configuration-invalid` | The plan does not include that engine. |
| 403, otherwise | `quota-exhausted` | The token grant is exhausted or expired. **No wait hint is set** — it does not clear on its own. |
| 404 | `configuration-invalid` | Nothing under that id. A jobId from another key reads exactly like one that never existed; a withdrawn voice answers 404, not 400. |
| 413 | `configuration-invalid` | Payload too large. |
| 422 | `configuration-invalid` | Moderation refused the text. Only reachable with AI Refine on. Rephrase; retrying unchanged fails again. |
| 429 | see below | Three different limiters. |
| 503 | `temporarily-unavailable` | No worker free for that engine. Transient. |
| other 5xx | `temporarily-unavailable` | Retry shortly. |
| anything else | `node-defect` | Unexpected response. |
VieNeu answers **403** when tokens run out — never 402. Retry logic keyed on 402
will never fire.
DNS, TLS, reset and timeout failures never reach that table; they become a
`temporarily-unavailable` error reading "Could not reach VieNeu while …", whose
description tells you to check that the Base URL includes `/api/v1`.
### Rate-limited versus quota-exhausted
A 429 can come from three places that want opposite responses. The node reads the
response headers to tell them apart — which is why it inspects the status itself
rather than letting n8n throw and lose them.
| Signal on the response | Which limiter | `failure.cause` | Wait hint | What to do |
| --- | --- | --- | --- | --- |
| `Retry-After` present | The app throttler, counted **per API key** | `rate-limited` | `retryAfterMs`, from the header | Wait the stated time and retry. Nothing is wrong with the balance. |
| No `Retry-After`, body matching `token limit reached` / `resets at` / `quota` | The **token quota** — a daily or weekly cap on the grant | `quota-exhausted` | `resetsAtEpochMs`, parsed out of the message text | Do not retry. Wait for the named reset, or top up. |
| No `Retry-After`, no such wording (nginx HTML) | The **nginx edge**, counted per source IP | `rate-limited` | none | Back off hard. On n8n Cloud that IP is shared with unrelated customers. |
Two of the three collapse onto `rate-limited`. Tell them apart by whether
`retryAfterMs` is set: present means the app throttler and a known wait; absent
means the edge and no wait hint at all.
The quota reset timestamp is parsed out of prose because it has to be: the
controller keeps only the message string, so the text is the last surviving trace
of the structured `resetAt` field.
Note that two different statuses map to `quota-exhausted`: the 429 above, which
carries a reset, and the 403 grant exhaustion, which carries none.
### Failures that are not HTTP failures
- **A failed job returns HTTP 200** with `status: "failed"`. Speech: Generate
raises it as an error; Job: Get returns it as data, so a Wait + IF loop can
see the terminal state and leave the loop.
- **A poll timeout** raises a 504-shaped error stating that the job is still
running and has **already been billed**, naming its `jobId`. Collect it later
with Job: Get. Keep Job Timeout under n8n's own execution timeout, or n8n kills
the run and the jobId goes with it.
- **An empty audio body**, or a submit that returns no `jobId`, raises a
502-shaped error.
- **A 200 with an empty voice list and an `error` field** — the catalogue read
failing — is raised as an error rather than rendered as an empty dropdown.
When the node is set to continue on failure instead of stopping, the failure
becomes item JSON: `error` with the redacted message, plus `jobId` **only when a
job was already submitted**. Nothing else is on that item — in particular, no
binary. A short-text failure, and an async failure that happened at submit, carry
`error` with no `jobId` at all.
## Example workflow: narrate long text, recover the ones that time out
Seven minutes of audio can outlast a poll window, and the job is billed at
submit. This workflow keeps those jobs instead of paying for nothing.
If you can receive an inbound HTTPS request, a
[webhook](../cloud-api/webhooks) and an n8n **Webhook** trigger node avoid the
whole problem — there is no poll window to outlast. Build this only when you
cannot.
1. **Manual Trigger** — "When clicking 'Execute workflow'". Replace it with
whatever produces your items; each item needs a `text` field.
2. **VieNeu** — Resource *Speech*, Operation *Generate*.
- **Text**: `={{ $json.text }}`
- **Engine**: pick one, so the voice list is scoped to it.
- **Voice**: *From List*, search for the voice you want.
- **Put Output File in Field**: `data`
- **Options → Format**: `WAV` (long text returns nothing else).
- **Options → Job Timeout (Seconds)**: `900`, below your n8n execution
timeout.
- Open the node's **Settings** tab and change **On Error** from *Stop
Workflow* to the option that continues using the node's **regular output**
— not the one that adds a separate error output connector, which would send
failed items down a branch the IF in step 3 never sees. Without this, one
timed-out job aborts the whole run and the jobId is lost.
3. **IF** — "Failed?". Condition: `={{ $json.error }}` *is not empty*. True means
something went wrong; false means the item has its audio on `data`.
4. **True branch → IF** — "Recoverable?". Condition: `={{ $json.jobId }}` *is not
empty*. True means the job was billed and is probably still running. False
means the failure happened before a job existed — a bad key, a wrong voice, a
429 — so there is nothing to collect: send it to an error output, a Slack
message, or wherever you want to see it. Do **not** merge it into step 8; it
has no binary and never will.
5. **Recoverable → Wait** — 60 seconds. Polling costs no tokens and no
rate-limit budget, but it is not free at the edge.
6. **Wait → VieNeu** — Resource *Job*, Operation *Get*.
- **Job ID**: `={{ $json.jobId }}`
- **Download Audio**: on (it is off by default on this operation)
- **Put Output File in Field**: `data`
7. **IF** — "Terminal?". Condition: `={{ $json.status }}` equals `completed`.
True → step 8. False → an IF on `={{ $json.status }}` equals `failed`, whose
true side ends the branch (a No Operation node) and whose false side loops
**back to the Wait node** in step 5. A job that outlasted a 900-second poll is
not reliably finished 60 seconds later, and without this loop the item falls
through with no binary attached.
8. Merge the completed items — those from step 3's false branch and those from
step 7's true branch — into whatever consumes the file: upload, email,
storage. Both carry the same binary property name, so downstream nodes do not
need to know which path an item took.
The same shape works item-by-item over a spreadsheet. Remember that each row is a
separate billed synthesis.
## Don't want to install the node?
Every operation is one HTTPS call, and n8n's built-in **HTTP Request** node makes
all of them. This path works on n8n Cloud, needs no build step, and needs nothing
from VieNeu but a key.
Two ready-made workflows ship with the package, under
`n8n-nodes-vieneu/templates/`:
| File | What it builds |
| --- | --- |
| `vieneu-speech-http-request.json` | Manual trigger → one HTTP Request to `POST /audio/speech` → MP3 on the `data` binary field. |
| `vieneu-long-text-http-request.json` | Submit to `POST /tts`, then Wait → poll → IF until the job is terminal, then download the presigned audio. |
Import either with *Workflows → Import from File*. Both attach the key through a
Bearer Auth credential, never a header typed into a node parameter. Note that
neither has been executed against a live n8n — they are validated structurally.
If you do not have the package, the recipes below build the same two workflows by
hand. For the underlying request contract, see the
[Cloud API overview](../cloud-api/overview).
Find a voice id first. Both calls below need no key. Since 2026-09-24 the cloud
API has one engine, V4, whose ids are display names (`Ngọc Lan`); the `vieneu-…`
slugs of the retired V3 catalogue no longer resolve and 400 on `voiceId`, so
refresh any id you saved before then.
```bash
# Which engine is the default? Look for "isDefault": true.
curl -s https://api.vieneu.io/api/v1/engines
# Then list that engine's voices only.
curl -s "https://api.vieneu.io/api/v1/voices?engine=v4" | head -c 400
```
Check that your key works, and hear the result, before wiring anything:
```bash
curl -X POST https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{"input":"Xin chào Việt Nam.","response_format":"mp3","speed":1}' \
--output speech.mp3
```
### Short text: one HTTP Request node
| Setting | Value |
| --- | --- |
| Method | `POST` |
| URL | `https://api.vieneu.io/api/v1/audio/speech` |
| Authentication | Generic Credential Type → **Bearer Auth**, holding your `vn_sk_…` key |
| Send Body | on, JSON |
| Response → Format | **File**, output property `data` |
Body, as an expression so the text can come from the incoming item:
```js
{{ JSON.stringify({
input: $json.text || 'Xin chào Việt Nam! Đây là giọng đọc tiếng Việt của VieNeu.',
voice: $json.voiceId || undefined,
response_format: 'mp3',
speed: 1
}) }}
```
Use a **Bearer Auth credential**, not an `Authorization` header typed into the
node. Credentials are referenced by id and are not part of the exported workflow
JSON — that is the difference between a workflow you can paste into an issue and
one that leaks your key when you do.
`response_format` accepts `mp3`, `wav`, `opus`, `pcm`, `ulaw` here. This route
blocks for the whole synthesis, so keep it for short text. A `clone_…` voice is
rejected here with a 400, same as through the node.
### Long text: submit, poll, download
Past roughly 500 characters, queue the job instead of holding a connection open.
Before you build the loop: if your n8n can receive an inbound HTTPS request, a
[webhook](../cloud-api/webhooks) replaces steps 2 to 5 with a single **Webhook**
trigger node. VieNeu POSTs a signed event the moment the job reaches a terminal
state, and `POST /tts` is the only route that produces those events. Poll only
when an inbound endpoint is not available to you.
The polling version, matching `vieneu-long-text-http-request.json`, is eight
nodes: a Manual Trigger and the seven below.
1. **HTTP Request — "Submit TTS job"**: `POST https://api.vieneu.io/api/v1/tts`,
Bearer Auth credential, JSON body:
```js
{{ JSON.stringify({
text: $json.text,
voiceId: $json.voiceId || undefined,
speed: 1
}) }}
```
Note the field names on this route: `text` and `voiceId`, not `input` and
`voice`. It takes no format parameter and always returns WAV. It responds with
a `jobId`, and the text is billed **here**, before the job is queued.
2. **Wait** — 3 seconds.
3. **HTTP Request — "Check job status"**:
`GET https://api.vieneu.io/api/v1/tts/{{ $json.jobId }}` (as an expression),
same Bearer Auth credential.
4. **IF — "Completed?"**: `={{ $json.status }}` equals `completed`. True →
download. False → step 5.
5. **IF — "Failed?"**: `={{ $json.status }}` equals `failed`. True → step 6.
False → back to the **Wait** node. A failed job answers HTTP **200**, so
branch on the body, never the status code.
6. **No Operation — "Job failed"**: ends the failed branch.
7. **HTTP Request — "Download audio"**: `GET {{ $json.audioUrl }}`, Response
Format **File**, output property `data`, and **no credential and no
authentication**. The URL is a presigned S3 link carrying its own signature —
S3 rejects the request outright if an `Authorization` header rides along.
Errors on this path arrive raw, without the node's classification. The rules
still hold: a 429 with `Retry-After` is the per-key throttler; a 429 whose
message mentions a token limit or a reset time is the quota; a 429 with neither
came from the edge and is counted per source IP. Running out of tokens entirely
is a 403.
---
# Installation
Source: https://docs.vieneu.io/docs/getting-started/installation
## Prerequisites
- **Python 3.10+**
- **eSpeak NG** — Required for phonemization
### Install eSpeak NG
```bash
# macOS
brew install espeak
# Ubuntu/Debian
sudo apt install espeak-ng
# Fedora/Amazon Linux
sudo dnf install espeak
# Windows
# Download .msi from https://github.com/espeak-ng/espeak-ng/releases
```
### Optional: NVIDIA GPU
For maximum speed via LMDeploy or GGUF GPU acceleration:
- NVIDIA Driver >= 570.65 (CUDA 12.8+)
- [NVIDIA GPU Computing Toolkit](https://developer.nvidia.com/cuda-downloads)
## Install from Source (Recommended)
```bash
git clone https://github.com/pnnbao97/VieNeu-TTS.git
cd VieNeu-TTS
```
### GPU Support (Default)
```bash
uv sync
```
### CPU-Only (Lightweight)
```bash
# Linux/macOS
cp pyproject.toml pyproject.toml.gpu
cp pyproject.toml.cpu pyproject.toml
uv sync
```
## Install as Python Package
```bash
# Windows (CPU optimized)
pip install vieneu --extra-index-url https://pnnbao97.github.io/llama-cpp-python-v0.3.16/cpu/
# macOS (Metal GPU accelerated)
pip install vieneu --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/metal/
# Linux / Generic
pip install vieneu
```
## Verify Installation
```python
from vieneu import Vieneu
tts = Vieneu()
audio = tts.infer(text="Xin chào")
tts.save(audio, "test.wav")
print("Installation successful!")
```
---
# Quick Start
Source: https://docs.vieneu.io/docs/getting-started/quick-start
## Web UI
The fastest way to try VieNeu-TTS:
```bash
uv run vieneu-web
```
Open `http://127.0.0.1:7860` — type text, pick a voice, click generate.
## Real-time Streaming
For ultra-low latency streaming (CPU optimized):
```bash
uv run vieneu-stream
```
Open `http://localhost:8001` — audio starts playing before the sentence finishes.
## Python SDK
### Basic Usage
```python
from vieneu import Vieneu
tts = Vieneu()
# Generate speech with default voice
audio = tts.infer(text="Xin chào, tôi là VieNeu.")
tts.save(audio, "output.wav")
```
### Voice Cloning
```python
audio = tts.infer(
text="Đây là giọng nói được clone.",
ref_audio="path/to/reference.wav",
ref_text="Transcript of the reference audio."
)
tts.save(audio, "cloned.wav")
```
### Using Preset Voices
```python
# List available voices
voices = tts.list_preset_voices()
for description, voice_id in voices:
print(f"{voice_id}: {description}")
# Use a specific voice
voice = tts.get_preset_voice("voice_name")
audio = tts.infer(text="Chào bạn!", voice=voice)
```
### Streaming
```python
for audio_chunk in tts.infer_stream(text="Một đoạn văn dài..."):
# Process each audio chunk as it's generated
play_audio(audio_chunk)
```
### Batch Processing
```python
texts = ["Câu một.", "Câu hai.", "Câu ba."]
audios = tts.infer_batch(texts)
for i, audio in enumerate(audios):
tts.save(audio, f"output_{i}.wav")
```
---
# System Requirements
Source: https://docs.vieneu.io/docs/getting-started/system-requirements
## Minimum Requirements
| Component | Requirement |
|-----------|------------|
| Python | 3.10+ |
| RAM | 2 GB (GGUF Q4) |
| Disk | ~500 MB (model auto-downloaded) |
| eSpeak NG | Required |
## Recommended (CPU)
| Component | Recommendation |
|-----------|---------------|
| CPU | Modern i5/i7/M1+ |
| RAM | 4 GB+ |
| Model | GGUF Q4 or Q8 |
Streaming latency: under 300ms first chunk on modern i3/i5.
## Recommended (GPU)
| Component | Recommendation |
|-----------|---------------|
| GPU | NVIDIA with 4GB+ VRAM |
| Driver | >= 570.65 (CUDA 12.8+) |
| Model | PyTorch 0.5B or 0.3B |
| Backend | LMDeploy for maximum speed |
## Intel Arc GPU
Supported via PyTorch XPU (tested on Arc B580, A770 on Windows):
```bash
run setup_xpu_uv.bat
run run_xpu.bat
```
Tip: Intel Arc has high memory bandwidth — keep batch size high and minimize characters per chunk.
## Model Sizes
| Model | Disk | RAM Usage |
|-------|------|-----------|
| 0.5B PyTorch | ~2 GB | ~3 GB |
| 0.3B PyTorch | ~1.2 GB | ~2 GB |
| 0.3B GGUF Q8 | ~350 MB | ~500 MB |
| 0.3B GGUF Q4 | ~200 MB | ~300 MB |
Models are cached at `~/.cache/huggingface/hub/` after first download.
---
# SDK Overview
Source: https://docs.vieneu.io/docs/sdk/overview
The `vieneu` package runs **VieNeu-TTS v3 Turbo** on your own machine. It defaults to the torch-free ONNX engine on CPU and switches to PyTorch on a CUDA GPU with no code change.
:::tip Where does this fit?
The SDK is the **on-device** path — free, open source (Apache 2.0), your hardware. If you would rather call a hosted API (including the proprietary **v4** engine with higher cloning fidelity), see the [Cloud API](/docs/cloud-api/overview).
:::
The pages after this one go deeper on one topic each:
- [Install & backends](/docs/sdk/standard-mode) — `pip install vieneu`, CPU vs GPU, precision, v3 Nano
- [GPU batching](/docs/sdk/fast-mode) — `infer_batch`, CUDA graphs, throughput numbers
- [Streaming](/docs/sdk/streaming) — `infer_stream`, concurrent streams
- [Voice cloning](/docs/sdk/voice-cloning) — `ref_audio`, `add_voice`, `denoise`
- [OpenAI-compatible server](/docs/sdk/remote-mode) — `/v1/audio/speech` from the repo or Docker, plus the legacy v2 `remote` mode
:::info Source
This section mirrors the **Using the Python SDK** part of the open-source [README](https://github.com/pnnbao97/VieNeu-TTS#readme) and is refreshed automatically (last sync 2026-09-16). If something here disagrees with the README, the README wins — [open an issue](https://github.com/pnnbao97/VieNeu-TTS/issues) there.
:::
The `vieneu` SDK **defaults to VieNeu-TTS v3 Turbo (48 kHz)**. The minimal install is **torch-free**: on CPU everything runs on **ONNX Runtime** (PyTorch is never imported), and on a CUDA machine it auto-switches to the PyTorch engine — where inference is **batched automatically** (same API, no code change).
## Quick Start
**CPU (default)** — torch-free, runs v3 Turbo via ONNX Runtime. Most users want this:
> ⚡**On CPU the backbone runs `fp32` by default** (maximum fidelity). Need more speed? Pass `Vieneu(precision="int8")` — ~1.6× faster and ~4× smaller, but it requires a CPU with VNNI (AVX-512 VNNI / AVX-VNNI); on older CPUs int8 can produce garbled audio. `precision` only affects the CPU/ONNX path; on GPU it's ignored (PyTorch).
>
> 🪶 **Still too slow, or deploying on a phone / ARM board?** Try **[VieNeu-TTS v3 Nano (preview)](#v3-nano)** — `Vieneu(mode="v3nano")`, ~3× faster than Turbo fp32 on CPU (RTF 0.11–0.22 on a desktop CPU), but **noticeably lower quality** (especially English / bilingual), 24 kHz, 11 preset voices + voice cloning. Details and caveats in the [v3 Nano section](#v3-nano) below.
```bash
pip install vieneu
```
**GPU (CUDA)** — only if you have an NVIDIA GPU. On Linux `pip install "vieneu[cuda]"` is enough (PyPI torch ships CUDA there); on Windows install the CUDA torch **first** as below.
> ℹ️ **How fast is the GPU path?** Since 3.7.0 every audio frame is **one CUDA
> graph** (acoustic decoder + sampling + repetition penalty + backbone step in a
> single replay — no `torch.compile`, no C++ toolchain needed). Measured on an
> RTX 3060: a 3.5 s sentence in **0.36 s**; a 2-chunk paragraph (19 s) in
> **1.4 s**; 16 chunks (154 s) in **2.8 s** (RTF 0.02) — previously 2.3 s /
> 8.7 s / 16.7 s. The first call for each batch size pays ~0.5 s to capture the
> graph (kept afterwards; servers can call `warm_fused()` at start-up).
> `VIENEU_FUSED_FRAME=0` restores the plain loop.
```bash
pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
pip install "transformers==4.57.6" # pinned — most stable transformers for the GPU SDK
pip install vieneu
```
```python
import time
from vieneu import Vieneu
# Default = v3 Turbo (48 kHz). GPU → PyTorch (auto-detected).
vieneu = Vieneu() # On a GPU machine you can still switch to ONNX/CPU if you prefer: Vieneu(backend="onnx")
# 1. Built-in voice by name — no reference clip needed
print("🔊 Generating speech...")
start_time = time.time()
audio = vieneu.infer("[cười] Trời ơi, cái giọng nó tự nhiên mà nó mượt mà dã man, nghe không khác gì người thật luôn. Giờ thì tha hồ mà quẩy content với cả kho giọng nói đa dạng, đủ mọi sắc thái biểu cảm. Mọi người bật loa lên rồi cùng trải nghiệm thử với mình nhé!", voice="Phạm Tuyên")
elapsed_time = time.time() - start_time
vieneu.save(audio, "output.wav")
print("✅ Saved to output.wav")
# Tính RTF (Real-Time Factor)
sample_rate = 48000
audio_duration = len(audio) / sample_rate
rtf = elapsed_time / audio_duration
print(f"\n⏱️ Thời gian xử lý: {elapsed_time:.3f}s")
print(f"🎵 Thời lượng audio: {audio_duration:.3f}s")
print(f"📊 RTF: {rtf:.4f} ({'nhanh hơn' if rtf < 1 else 'chậm hơn'} real-time {1/rtf:.2f}x)" if rtf > 0 else "")
# List the built-in voices
voices = vieneu.list_preset_voices()
print(f"\n🎙️ {len(voices)} built-in voices available:")
for label, voice_id in voices:
print(f" - {label} ({voice_id})")
# 2. ⚡ Batch on GPU: infer_batch() runs many texts in ONE batched forward — same API.
# On a CUDA GPU the chunks from every text share each forward step (big throughput
# win). On CPU it still WORKS (no error) — just sequentially, so there's no batch
# gain. Batch caps at max_batch_size (default 32; tune via Vieneu(max_batch_size=64)
# or infer_batch(..., batch_size=64), or batch_size=1 to disable). A single long
# infer() also auto-batches its own chunks. For real-time use, infer_stream() is the
# streaming twin (GPU: 16 concurrent streams — see "Streaming" below). Uncomment to
# try (GPU recommended):
#
# import time
# texts = [
# "Chào cả nhà, hôm nay mình sẽ hướng dẫn các bạn cách cài đặt và sử dụng bộ giọng đọc mới.",
# "Giọng nghe cực kỳ tự nhiên và truyền cảm, lại có thể chuyển đổi biểu cảm một cách linh hoạt.",
# "Nếu thấy hữu ích, các bạn nhớ để lại một lượt thích và chia sẻ video này cho mọi người nhé!",
# ] * 10 # 30 texts — enough to fill the batch and really show the GPU throughput win
# t0 = time.time()
# audios = vieneu.infer_batch(texts, voice="Minh Quân Pro")
# elapsed = time.time() - t0
# total_audio = sum(len(a) for a in audios) / 48_000
# print(f"⚡ {len(texts)} texts | audio {total_audio:.1f}s | wall {elapsed:.1f}s | RTF {elapsed/total_audio:.3f}")
# for i, a in enumerate(audios):
# vieneu.save(a, f"batch_{i}.wav")
```
### Streaming (real-time) 🔊
> v3 Turbo streams **frame by frame** on both backends. **GPU** (PyTorch): first audio in **~115 ms** and **16 concurrent streams** on one RTX 3060 (continuous batching — one CUDA graph serves every `infer_stream` call, each keeping RTF ≈ 0.5–0.6). **CPU** (ONNX): first audio in ~140 ms (int8) / ~300 ms (fp32), one stream (two with int8). Just iterate `infer_stream`:
```python
from vieneu import Vieneu
vieneu = Vieneu() # GPU → PyTorch + stream scheduler; no GPU → ONNX/CPU
for chunk in vieneu.infer_stream("Xin chào các bạn!", voice="Mai Anh"):
play(chunk) # np.float32 @ 48 kHz — play/write as it arrives
```
Calling `infer_stream` from many threads at once is the intended way to serve many listeners on a GPU (`Vieneu(max_streams=16)` sets the ceiling).
An **OpenAI-compatible streaming API** (`POST /v1/audio/speech`, `pcm`/`wav`, chunked or SSE — works with the OpenAI SDK, Pipecat, LiveKit, …) is in [`apps/openai_speech.py`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/apps/openai_speech.py):
```bash
# Pick ONE of these — all serve http://localhost:8000/v1/audio/speech
uv run python -m apps.openai_speech # from the repo (auto-detects GPU/CPU)
docker compose -f docker/docker-compose.yml --profile api-gpu up # or: Docker, GPU
docker compose -f docker/docker-compose.yml --profile api-cpu up # or: Docker, CPU only
```
📊 **[docs/streaming.md](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/streaming.md)** — every measurement on an RTX 3060 (TTFA / RTF / streams vs `max_streams`), estimates for smaller GPUs, and the CPU numbers. The older browser demo is still at [`apps/web_stream.py`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/apps/web_stream.py).
### Available Voices
The v3 Turbo engine includes **25 preset voices** covering **3 regions** (North, Central, South) with diverse genders and speaking characters. `list_preset_voices()` (and the Web UI / API voice lists) show them in this order:
- ⭐ **Editors' picks** — the 10 we recommend starting with, hand-selected for naturalness and stability: **Adam bựa, Trúc Ly, Anh Khôi, Mai Anh, Minh Quân Pro** *(default; `"Minh Quân"` still works as an alias)*, **Thùy Dung, Thiền Tâm Đức, Ngọc Huyền, Quang Sơn, Ngọc Trân**
- **Northern (Bắc)**: Minh Đức, Phạm Tuyên, Xuân Vĩnh, Thanh Bình, Ngọc Linh, Đoan Trang, Quỳnh Anh, Mạnh Dũng (+ picks above)
- **Central (Trung)**: Quang Sơn, Ngọc Trân
- **Southern (Nam)**: Adam, Thái Sơn, Thục Đoan, Minh Triết, Mỹ Duyên, Đức Trí, Kim Thanh (+ Thùy Dung)
## Reading style — **deprecated** ⚠️
:::warning
**`style` is deprecated on v3 Turbo and has no effect.** The reading style is already
baked into the reference itself (the speaker embedding + reference codes of the preset
voice or of your cloned clip), so every generation follows the reference and comes out
in its natural reading style.
The `style` argument is **still accepted** by `infer`, `infer_stream`, `infer_batch`
and `add_voice` so existing code keeps running — whatever you pass (`"tin_tuc"`,
`"doc_truyen"`, …) is simply ignored. New code should just omit it.
:::
```python
# Old code — still runs, but `style` is ignored
audio = vieneu.infer("Bản tin sáng nay.", voice="Minh Quân Pro", style="tin_tuc")
# New code — pick the reading character through the voice / reference clip instead
audio = vieneu.infer("Bản tin sáng nay.", voice="Minh Quân Pro")
```
## Emotion cues (experimental)
Inline tags are supported anywhere in the text: `[cười]` (chuckle), `[thở dài]` (sigh), `[hắng giọng]` (clear throat).
```python
audio = vieneu.infer("Nghe hay quá đi [cười]. Để mình nói tiếp [hắng giọng].", voice="Minh Quân Pro")
```
## Voice cloning
Clone any voice from a short reference clip. The clip is cleaned up automatically
(background noise removed, and trimmed to ≤ 8 seconds) before cloning — keep
`denoise=True` unless your clip is already clean.
```python
audio = vieneu.infer(
"Đây là giọng được nhân bản tức thì.",
ref_audio="my_voice.wav", # a 3–8s reference clip
denoise=True, # default; set False if the clip is already clean
)
vieneu.save(audio, "cloned.wav")
```
## Save & reuse a cloned voice
Register a reference once with `add_voice`, then use it by name like a built-in voice.
```python
# Enroll a voice (denoises + extracts the speaker profile once)
vieneu.add_voice("Giọng của tôi", "my_voice.wav")
# Now reuse it anywhere, including the conversation mode
audio = vieneu.infer("Câu này dùng giọng đã lưu.", voice="Giọng của tôi")
# Persist your voices so they load next session
vieneu.save_voices() # writes to the default voices file
# vieneu.remove_voice("Giọng của tôi")
# Add a voice you already cleaned yourself → skip denoising
vieneu.add_voice("Giọng sạch", "already_clean.wav", denoise=False)
```
## Clean up a clip on its own
Get the denoised audio without synthesizing anything (e.g. to inspect or store it):
```python
wav, sr = vieneu.denoise("noisy.wav", out_path="clean.wav") # 44.1 kHz mono
```
> **Note:** `denoise`, `add_voice`, and voice cloning work on every backend — the
> torch-free CPU/ONNX install included (the whole cloning pipeline runs on
> onnxruntime + soxr + kaldi-native-fbank). **v3 Nano** below clones the same way (its cloning graphs are fetched on first use).
## v3 Nano (preview) — for edge devices / weak CPUs only 🪶 {#v3-nano}
:::warning
**v3 Turbo remains the default and the recommended model.** Use v3 Nano only when Turbo is
too slow on your hardware (old laptops, mini PCs, ARM boards, CPUs without AVX-512/VNNI where
the int8 Turbo build produces garbled audio). Nano is a 48M-parameter flow-matching model
(ONNX, CPU, torch-free) and it **trades quality for speed**:
- **Lower quality than v3 Turbo — most noticeably on English and code-switched (En-Vi) text.**
Vietnamese is close; English words come out with a Vietnamese accent and are less stable.
- **24 kHz** output (Turbo: 48 kHz).
- **11 preset voices + voice cloning** (`ref_audio`, `add_voice`, `encode_reference` work like Turbo; the three cloning graphs, ~110 MB, download on first use).
- **No frame-level streaming** — `infer_stream` yields one finished chunk at a time.
:::
Measured on the same desktop CPU (12th-gen Intel i7, 6 ONNX Runtime threads, ~9 s of speech):
| Engine | RTF ↓ | Sample rate | Load time |
|---|---|---|---|
| v3 Turbo ONNX fp32 (default on CPU) | 0.62 | 48 kHz | ~19 s |
| v3 Turbo ONNX int8 | 0.37 | 48 kHz | ~14 s |
| **v3 Nano, 16 steps, cfg 3** (default) | **0.22** | 24 kHz | ~3 s |
| **v3 Nano, 8 steps, sway −1** | **0.11** | 24 kHz | ~3 s |
RTF = compute time ÷ audio duration (lower is faster; 0.22 = 4.5× faster than real time). The ratio carries over to slower machines: expect Nano to be roughly **1.7× faster than Turbo int8** and **~3× faster than Turbo fp32**, with a 282 MB download instead of Turbo's.
```python
from vieneu import Vieneu
tts = Vieneu(mode="v3nano") # ONNX, CPU, torch-free
audio = tts.infer("Xin chào, mình là giọng đọc của VieNeu Nano.", voice="Minh Quân")
tts.save(audio, "nano.wav") # 24 kHz
tts.list_preset_voices() # Adam, Ái Hân, Mỹ Duyên, Đức Trí, Hữu Quân, Xuân Tiên, Mai Anh, Trúc Ly, Anh Khôi, Minh Quân, Mạnh Dũng
audio = tts.infer("Bản nhanh cho máy rất yếu.", voice="Ái Hân", steps=8, sway=-1) # ~2× faster
```
Knobs: `steps` (Euler steps, 16 default; 8 ≈ 2× faster, slightly rougher — pair with `sway=-1`),
`cfg` (classifier-free guidance, 3.0 default; `cfg=0` halves compute but hurts intelligibility),
`speed`, `seed`, `threads`. Emotion cues `[cười]` `[thở dài]` `[hắng giọng]` work as on Turbo.
---
# Install & backends
Source: https://docs.vieneu.io/docs/sdk/standard-mode
One package, two engines. `Vieneu()` picks the engine from your hardware; the API is the same on both.
| You have | Engine | Install | Notes |
|---|---|---|---|
| CPU only, macOS | ONNX Runtime, **torch-free** | `pip install vieneu` | 48 kHz v3 Turbo, cloning and emotion cues included. PyTorch is never installed. |
| NVIDIA GPU | PyTorch (CUDA) | `pip install "vieneu[cuda]"` | Batched automatically; every frame is one CUDA graph since 3.7.0. |
| Weak CPU / ARM board | ONNX, **v3 Nano** | `pip install vieneu` + `Vieneu(mode="v3nano")` | ~3× faster than Turbo fp32, 24 kHz, lower quality (esp. English). |
## CPU (default)
```bash
pip install vieneu
```
```python
from vieneu import Vieneu
tts = Vieneu() # v3 Turbo, ONNX, fp32
audio = tts.infer("Xin chào bạn", voice="Minh Quân Pro")
tts.save(audio, "output.wav") # 48 kHz WAV
```
- **Precision.** `fp32` is the default for maximum fidelity. `Vieneu(precision="int8")` is ~1.6× faster and ~4× smaller, but needs a CPU with VNNI (AVX-512 VNNI / AVX-VNNI). On older CPUs int8 can produce garbled audio; if that happens, go back to fp32 or try v3 Nano.
- **Fastest CPU install.** From a checkout of the repo, `uv sync` reproduces the locked environment that pins the optimised ONNX Runtime build. It is measurably faster than a plain `pip install`.
- **Apple Silicon.** Use the CPU/ONNX path. It is faster than the MPS/PyTorch build for v3 Turbo.
## GPU (CUDA)
On Linux the PyPI torch wheel already ships CUDA:
```bash
pip install "vieneu[cuda]"
```
On Windows install the CUDA torch **first**, then a pinned transformers, then vieneu:
```bash
pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
pip install "transformers==4.57.6"
pip install vieneu
```
```python
tts = Vieneu() # CUDA detected → PyTorch engine
tts = Vieneu(backend="onnx") # force the CPU engine on a GPU machine
```
`precision` only applies to the ONNX path and is ignored on GPU. Throughput numbers and batching are on the [GPU batching](/docs/sdk/fast-mode) page.
## v3 Nano (preview)
A 48M-parameter flow-matching model for hardware where Turbo is too slow. Same cloning API, 11 preset voices, 24 kHz output, no frame-level streaming (chunks arrive whole).
```python
tts = Vieneu(mode="v3nano")
audio = tts.infer("Bản nhanh cho máy rất yếu.", voice="Ái Hân", steps=8, sway=-1)
```
Measured on a 12th-gen i7 desktop, 6 ONNX threads (RTF = compute ÷ audio duration, lower is faster):
| Engine | RTF | Sample rate | Load |
|---|---|---|---|
| v3 Turbo fp32 (CPU default) | 0.62 | 48 kHz | ~19 s |
| v3 Turbo int8 | 0.37 | 48 kHz | ~14 s |
| v3 Nano, 16 steps (default) | 0.22 | 24 kHz | ~3 s |
| v3 Nano, 8 steps, sway −1 | 0.11 | 24 kHz | ~3 s |
Knobs: `steps` (16 default, 8 ≈ 2× faster), `cfg` (3.0 default, 0 halves compute but hurts intelligibility), `speed`, `seed`, `threads`.
## Models and cache
Backbone and codec weights download from Hugging Face on first use and are cached under `~/.cache/huggingface/hub/`. Turbo's cloning pipeline runs on onnxruntime + soxr + kaldi-native-fbank; Nano fetches its three cloning graphs (~110 MB) the first time you clone.
## Legacy backends (v1 / v2)
The GGUF (llama-cpp) and LMDeploy backends serve **VieNeu-TTS v1/v2** only and are no longer updated. They live behind `uv sync --group gpu` in the repo and `pip install "vieneu[legacy]"`. New projects should stay on v3 Turbo.
---
# GPU batching
Source: https://docs.vieneu.io/docs/sdk/fast-mode
On a CUDA GPU v3 Turbo batches automatically. A long `infer()` batches its own chunks; `infer_batch()` batches many texts in one forward pass. Same API, no code change.
:::note Renamed page
This page used to describe the LMDeploy "fast mode" for VieNeu-TTS v2. That backend is legacy now; see the bottom of the page.
:::
## infer_batch
```python
from vieneu import Vieneu
tts = Vieneu() # CUDA → PyTorch engine
texts = [
"Chào cả nhà, hôm nay mình sẽ hướng dẫn các bạn cách cài đặt bộ giọng đọc mới.",
"Giọng nghe cực kỳ tự nhiên và truyền cảm.",
"Nếu thấy hữu ích, nhớ để lại một lượt thích nhé!",
] * 10
audios = tts.infer_batch(texts, voice="Minh Quân Pro")
for i, a in enumerate(audios):
tts.save(a, f"batch_{i}.wav")
```
- Chunks from every text share each forward step, which is where the throughput win comes from.
- Batch size caps at `max_batch_size` (default 32). Tune with `Vieneu(max_batch_size=64)` or `infer_batch(..., batch_size=64)`; `batch_size=1` disables batching.
- On CPU `infer_batch` still works, just sequentially, so there is no gain.
- For real-time playback use [`infer_stream`](/docs/sdk/streaming) instead. It is the streaming twin and serves 16 concurrent listeners on one GPU.
## CUDA graphs (3.7.0+)
Every audio frame is a single CUDA graph replay: acoustic decoder, sampling, repetition penalty and the backbone step in one shot. No `torch.compile`, no C++ toolchain.
Measured on an RTX 3060:
| Input | Audio | Wall time | Before 3.7.0 |
|---|---|---|---|
| One sentence | 3.5 s | 0.36 s | 2.3 s |
| Two-chunk paragraph | 19 s | 1.4 s | 8.7 s |
| 16 chunks | 154 s | 2.8 s (RTF 0.02) | 16.7 s |
The first call for each batch size pays ~0.5 s to capture the graph, then keeps it. Servers can call `tts.warm_fused()` at start-up. `VIENEU_FUSED_FRAME=0` restores the plain loop if you need to debug.
## Requirements
- NVIDIA GPU. An RTX 3060 (12 GB) gives the numbers above; ~6 GB is enough for inference.
- Install per [Install & backends](/docs/sdk/standard-mode#gpu-cuda): CUDA torch 2.8.0 first on Windows, `transformers==4.57.6`, then `vieneu`.
## Fine-tuned models
A LoRA-merged v3 Turbo keeps the full API, batching included:
```python
tts = Vieneu(mode="v3turbo", backbone_repo="finetune/output/my_voice/merged")
```
See [Fine-tuning](/docs/advanced/fine-tuning) and the repo's [`finetune/README.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/finetune/README.md).
## Legacy: LMDeploy "fast" mode (v2 only)
`Vieneu(mode="fast")` loads VieNeu-TTS v1/v2 through LMDeploy. It is kept for existing deployments (`pip install "vieneu[legacy]"`, or `uv sync --group gpu` in the repo) and receives no updates. v3 Turbo on PyTorch is faster and needs no extra runtime.
---
# OpenAI-compatible server
Source: https://docs.vieneu.io/docs/sdk/remote-mode
`apps/openai_speech.py` in the repo serves `POST /v1/audio/speech` exactly like OpenAI's TTS endpoint (`pcm`/`wav`, chunked body or SSE). The OpenAI SDK, Pipecat, LiveKit Agents and the Vercel AI SDK work by changing `base_url`.
## Start the server
Pick one; all listen on `http://localhost:8000`:
```bash
uv run python -m apps.openai_speech # from a repo checkout, auto-detects GPU/CPU
docker compose -f docker/docker-compose.yml --profile api-gpu up # Docker, GPU
docker compose -f docker/docker-compose.yml --profile api-cpu up # Docker, CPU only (torch-free)
```
Measure time-to-first-audio and RTF on your own machine:
```bash
uv run python examples/openai_speech_client.py --bench 8
```
## Call it
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")
with client.audio.speech.with_streaming_response.create(
model="vieneu-v3-turbo",
voice="Mai Anh",
input="Xin chào! Đây là chế độ streaming của VieNeu.",
response_format="pcm",
) as r:
for chunk in r.iter_bytes(4096): # s16le 48 kHz mono, as it is generated
play(chunk)
```
## Endpoints
| Method | Path | Purpose |
|---|---|---|
| `POST` | `/v1/audio/speech` | Synthesize; `response_format` `pcm` or `wav`, streamed |
| `GET` | `/v1/models` | Model list |
| `GET` | `/v1/voices` | Preset and enrolled voices |
| `POST` | `/v1/voices` | Clone from an uploaded clip |
| `GET` | `/health` | Liveness |
Concurrency is capped per backend with `VIENEU_MAX_STREAMS` (default 16 on GPU, 1 on CPU). Requests beyond the cap wait in a small queue, then get `429`. First audio arrives in ~115 ms with 16 streams on an RTX 3060; ~140–300 ms and 1–2 streams on CPU. Full numbers: [`docs/streaming.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/streaming.md).
## Web UI in Docker
```bash
docker compose -f docker/docker-compose.yml --profile gpu up # or --profile cpu → http://localhost:7860
```
Production images and builds: the repo's [`docs/Deploy.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/Deploy.md) and our [Docker](/docs/deployment/docker) page.
## Hosted instead of self-hosted
If you do not want to run a GPU, the [VieNeu Cloud API](/docs/cloud-api/openai-compatible) exposes the same OpenAI-compatible shape at `api.vieneu.io`, including the v4 engine.
## Legacy: v2 `remote` mode (deprecated)
:::warning
The LMDeploy server on port 23333 and `Vieneu(mode="remote")` only work with **VieNeu-TTS v2**, which is no longer updated. They are kept for existing deployments. For v3 Turbo use the streaming server above.
:::
```bash
docker run --gpus all -p 23333:23333 -v huggingface_cache:/root/.cache/huggingface pnnbao/vieneu-tts:latest --tunnel
pip install "vieneu[legacy]"
```
```python
from vieneu import Vieneu
tts = Vieneu(mode="remote", api_base="http://your-server-ip:23333/v1",
model_name="pnnbao-ump/VieNeu-TTS-v2", emotion="natural")
audio = tts.infer(text="Chào bạn!")
tts.save(audio, "remote_output.wav")
```
Fine-tuned v3 Turbo models are not served by that container; load them with the SDK (`Vieneu(mode="v3turbo", backbone_repo=...)`). See [Remote server](/docs/deployment/remote-server) for the old Docker flags.
---
# Streaming
Source: https://docs.vieneu.io/docs/sdk/streaming
v3 Turbo streams **frame by frame** on both engines. Iterate `infer_stream` and play or write each chunk as it arrives.
```python
from vieneu import Vieneu
tts = Vieneu() # GPU → PyTorch + stream scheduler; CPU → ONNX
for chunk in tts.infer_stream("Xin chào các bạn!", voice="Mai Anh"):
play(chunk) # np.float32 @ 48 kHz
```
## Latency and concurrency
| Engine | First audio | Concurrent streams |
|---|---|---|
| GPU (PyTorch), RTX 3060 | ~115 ms | 16 (32 max), each at RTF ≈ 0.5–0.6 |
| CPU (ONNX) int8 | ~140 ms | 2 |
| CPU (ONNX) fp32 | ~300 ms | 1 |
On a GPU the streams share one CUDA graph through continuous batching. Calling `infer_stream` from many threads at once is the intended way to serve many listeners; `Vieneu(max_streams=16)` sets the ceiling. The first request after the GPU has idled pays an extra 100–300 ms until it clocks up.
Every measurement (TTFA and RTF against `max_streams`, estimates for smaller GPUs, CPU numbers) is in the repo's [`docs/streaming.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/streaming.md).
## Chunk format
Each chunk is a `numpy.float32` array at 48 kHz, mono. Convert to 16-bit PCM for most audio devices:
```python
import numpy as np
pcm16 = (np.clip(chunk, -1, 1) * 32767).astype(np.int16).tobytes()
```
## Serving over HTTP
To expose streaming to other processes or languages, run the OpenAI-compatible server from the repo. It streams `pcm`/`wav` as chunked body or SSE from `POST /v1/audio/speech`, so the OpenAI SDK, Pipecat and LiveKit work by changing `base_url`. See [OpenAI-compatible server](/docs/sdk/remote-mode).
## v3 Nano
Nano has no frame-level streaming: `infer_stream` yields one finished chunk at a time. Use Turbo when time-to-first-audio matters.
## Hosted alternative
The same streaming shape is available without a GPU from the [Cloud API streaming endpoint](/docs/cloud-api/streaming).
---
# Voice cloning
Source: https://docs.vieneu.io/docs/sdk/voice-cloning
Clone any voice from a **3–8 second** clip. No transcript, no fine-tuning. The clip is denoised and trimmed to ≤ 8 s automatically before cloning.
```python
from vieneu import Vieneu
tts = Vieneu()
audio = tts.infer(
"Đây là giọng được nhân bản tức thì.",
ref_audio="my_voice.wav", # 3–8 s reference clip
denoise=True, # default; False if the clip is already clean
)
tts.save(audio, "cloned.wav")
```
:::info No `ref_text` on v3
v1/v2 needed the exact transcript of the reference clip (`ref_text`). v3 Turbo extracts a speaker embedding and reference codes instead, so a transcript is not required.
:::
## Enrol once, reuse by name
```python
tts.add_voice("Giọng của tôi", "my_voice.wav") # denoise + extract the profile once
audio = tts.infer("Câu này dùng giọng đã lưu.", voice="Giọng của tôi")
tts.save_voices() # persist to the default voices file
# tts.remove_voice("Giọng của tôi")
tts.add_voice("Giọng sạch", "already_clean.wav", denoise=False)
```
Enrolled voices work everywhere a preset does, including `infer_batch`, `infer_stream` and the conversation mode.
## Denoise on its own
```python
wav, sr = tts.denoise("noisy.wav", out_path="clean.wav") # 44.1 kHz mono
```
## Reading style follows the reference
The `style` argument (`tin_tuc`, `doc_truyen`, …) is **deprecated and ignored** on v3 Turbo. The reading character is baked into the reference: pick a preset voice or a clip that already reads the way you want. Passing `style` still runs for backwards compatibility.
## Emotion cues (experimental)
Inline tags work with cloned voices too: `[cười]` (chuckle), `[thở dài]` (sigh), `[hắng giọng]` (clear throat).
```python
audio = tts.infer("Nghe hay quá đi [cười]. Để mình nói tiếp [hắng giọng].", voice="Giọng của tôi")
```
## Backends
`denoise`, `add_voice` and cloning work on every backend, including the torch-free CPU/ONNX install (the pipeline runs on onnxruntime + soxr + kaldi-native-fbank). **v3 Nano** clones the same way; its three cloning graphs (~110 MB) download on first use.
## Tips for a good reference
- 3–8 s of one speaker, no music, no second voice.
- Natural, continuous speech beats isolated words.
- Keep `denoise=True` unless you cleaned the clip yourself.
- Want a tighter match or a specific reading style? [Fine-tune with LoRA](/docs/advanced/fine-tuning) on 10–30 minutes of audio.
## Higher fidelity: v4 on the Cloud API
The proprietary **VieNeu-TTS v4** reproduces a reference with near-original speaker similarity. It is not open source and is available only through the [Cloud API](/docs/cloud-api/overview).
---
# Vieneu() Factory
Source: https://docs.vieneu.io/docs/api/factory
The main entry point for creating a VieNeu-TTS instance.
## Signature
```python
Vieneu(mode="standard", **kwargs)
```
## Parameters
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `mode` | `str` | `"standard"` | Backend mode: `"standard"`, `"fast"`, `"remote"`, `"xpu"` |
### Standard Mode kwargs
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `backbone_repo` | `str` | `"pnnbao-ump/VieNeu-TTS-0.3B-q4-gguf"` | HuggingFace repo or local path |
| `backbone_device` | `str` | `"cpu"` | `"cpu"`, `"cuda"`, `"mps"` |
| `codec_repo` | `str` | `"neuphonic/distill-neucodec"` | Codec model repo |
| `codec_device` | `str` | `"cpu"` | Device for codec |
| `hf_token` | `str` | `None` | HuggingFace token for private models |
### Remote Mode kwargs
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `api_base` | `str` | required | Server URL (e.g., `"http://host:23333/v1"`) |
| `model_name` | `str` | required | Model ID on the server |
## Returns
An instance of `BaseVieneuTTS` (subclass depends on `mode`).
## Examples
```python
# Default: GGUF Q4 on CPU
tts = Vieneu()
# PyTorch on GPU
tts = Vieneu(backbone_repo="pnnbao-ump/VieNeu-TTS-0.3B", backbone_device="cuda")
# LMDeploy fast mode
tts = Vieneu(mode="fast", backbone_repo="pnnbao-ump/VieNeu-TTS")
# Remote client
tts = Vieneu(mode="remote", api_base="http://server:23333/v1", model_name="pnnbao-ump/VieNeu-TTS")
```
---
# Inference Methods
Source: https://docs.vieneu.io/docs/api/infer
## `infer()`
```python
audio = tts.infer(
text: str,
ref_audio: str = None,
ref_codes: Tensor = None,
ref_text: str = None,
voice: dict = None,
max_chars: int = 256,
silence_p: float = 0.15,
crossfade_p: float = 0.0,
temperature: float = 1.0,
top_k: int = 50,
skip_normalize: bool = False,
)
```
### Parameters
| Parameter | Type | Description |
|-----------|------|-------------|
| `text` | `str` | Text to synthesize |
| `ref_audio` | `str` | Path to reference audio for voice cloning |
| `ref_codes` | `Tensor` | Pre-encoded reference codes |
| `ref_text` | `str` | Transcript of reference audio |
| `voice` | `dict` | Preset voice dict from `get_preset_voice()` |
| `max_chars` | `int` | Max characters per chunk (default 256) |
| `silence_p` | `float` | Silence duration between chunks in seconds |
| `crossfade_p` | `float` | Crossfade duration between chunks |
| `temperature` | `float` | Sampling temperature |
| `top_k` | `int` | Top-k sampling |
| `skip_normalize` | `bool` | Skip text normalization |
### Returns
`numpy.ndarray` — Audio waveform at 24 kHz.
### Voice Priority
1. `voice` dict (from preset)
2. `ref_audio` + `ref_text`
3. `ref_codes` + `ref_text`
4. Default preset voice
---
## `infer_batch()`
```python
audios = tts.infer_batch(texts: List[str], ...)
```
Returns `List[numpy.ndarray]`. PyTorch mode uses true batch generation; GGUF processes sequentially.
---
## `infer_stream()`
```python
for chunk in tts.infer_stream(text: str, ...):
play_audio(chunk)
```
Yields `numpy.ndarray` chunks (GGUF only).
---
## `save()`
```python
tts.save(audio: numpy.ndarray, output_path: str)
```
---
## `encode_reference()`
```python
codes = tts.encode_reference(ref_audio_path: str)
# Returns: torch.Tensor
```
---
## `close()`
```python
tts.close()
# Or use context manager:
with Vieneu() as tts:
audio = tts.infer(text="...")
```
---
# Voice Management
Source: https://docs.vieneu.io/docs/api/voice-management
## Preset Voices
### `list_preset_voices()`
```python
voices = tts.list_preset_voices()
# Returns: List[tuple[str, str]] → [(description, voice_id), ...]
```
### `get_preset_voice()`
```python
voice = tts.get_preset_voice(voice_name: str = None)
# Returns: dict → {"codes": Tensor, "text": str}
```
### Using a Preset Voice
```python
voices = tts.list_preset_voices()
voice = tts.get_preset_voice("bac_si_tuyen")
audio = tts.infer(text="Chào bạn!", voice=voice)
```
## LoRA Adapters
### `load_lora_adapter()`
```python
success = tts.load_lora_adapter(
lora_repo_id: str,
hf_token: str = None,
)
```
### `unload_lora_adapter()`
```python
success = tts.unload_lora_adapter()
```
## voices.json Format
```json
{
"default_voice": "voice_name",
"presets": {
"voice_name": {
"description": "Description of the voice",
"text": "Transcript of the reference audio",
"codes": [42, 17, 89, 55, ...]
}
}
}
```
---
# Docker Deployment
Source: https://docs.vieneu.io/docs/deployment/docker
## Requirements
- Docker
- [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html)
## Docker Compose
```bash
docker compose -f docker/docker-compose.yml --profile gpu up
```
## Custom Docker Run
```bash
docker run --gpus all -p 7860:7860 pnnbao/vieneu-tts:latest
```
Access the Web UI at `http://localhost:7860`.
:::note
Docker deployment currently supports **GPU only**. For CPU, install from source.
:::
---
# Remote Server
Source: https://docs.vieneu.io/docs/deployment/remote-server
Deploy VieNeu-TTS as a high-performance API server powered by LMDeploy.
## Quick Start
```bash
docker run --gpus all -p 23333:23333 pnnbao/vieneu-tts:serve
```
## With Public Tunnel
```bash
docker run --gpus all -p 23333:23333 pnnbao/vieneu-tts:serve --tunnel
```
## Run 0.3B Model (Faster)
```bash
docker run --gpus all pnnbao/vieneu-tts:serve --model pnnbao-ump/VieNeu-TTS-0.3B --tunnel
```
## Serve a Fine-tuned Model
```bash
docker run --gpus all \
-v $(pwd)/finetune/output:/workspace/models \
pnnbao/vieneu-tts:serve \
--model /workspace/models/merged_model --tunnel
```
## Connecting Clients
See [Remote Mode](/docs/sdk/remote-mode) for client SDK usage.
---
# Custom Models
Source: https://docs.vieneu.io/docs/advanced/custom-models
## Via SDK
```python
from vieneu import Vieneu
# Load from HuggingFace
tts = Vieneu(backbone_repo="your-username/your-model")
# Load from local path
tts = Vieneu(backbone_repo="/path/to/your/model")
```
## Via Web UI
The Web UI provides a model selector where you can enter any HuggingFace repo ID or local path.
## LoRA Adapters
```python
tts = Vieneu(
backbone_repo="pnnbao-ump/VieNeu-TTS",
backbone_device="cuda",
)
tts.load_lora_adapter("your-username/your-lora-adapter")
audio = tts.infer(text="Chào bạn!")
tts.unload_lora_adapter()
```
:::note
LoRA adapters require PyTorch backbone. Not supported with GGUF models.
:::
---
# Fine-tuning
Source: https://docs.vieneu.io/docs/advanced/fine-tuning
Train VieNeu-TTS on your own voice or custom datasets using LoRA.
## Quick Start
```bash
cd finetune
python train.py --config config.yaml
```
## Google Colab
Use the provided notebook: `finetune/finetune_VieNeu-TTS.ipynb`
## Workflow
1. **Prepare data** — Audio files + transcripts
2. **Configure** — Edit training config (LoRA rank, learning rate, etc.)
3. **Train** — Run `train.py`
4. **Merge** — Merge LoRA weights into base model (optional)
5. **Use** — Load via `load_lora_adapter()` or serve merged model
## Documentation
See the detailed guide at [`finetune/README.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/finetune/README.md).
---
# Model Overview
Source: https://docs.vieneu.io/docs/advanced/model-overview
## Available Models
| Model | Format | Device | Quality | Speed |
|-------|--------|--------|---------|-------|
| VieNeu-TTS (0.5B) | PyTorch | GPU/CPU | Best | Very Fast (LMDeploy) |
| VieNeu-TTS-0.3B | PyTorch | GPU/CPU | Great | Ultra Fast (2x) |
| VieNeu-TTS Q8 GGUF | GGUF | CPU/GPU | Great | Fast |
| VieNeu-TTS Q4 GGUF | GGUF | CPU/GPU | Good | Very Fast |
| VieNeu-TTS-0.3B Q8 GGUF | GGUF | CPU/GPU | Great | Ultra Fast (1.5x) |
| VieNeu-TTS-0.3B Q4 GGUF | GGUF | CPU/GPU | Good | Extreme (2x) |
## Architecture
- **0.5B** — Fine-tuned from NeuTTS Air architecture. Maximum stability and quality.
- **0.3B** — Trained from scratch on VieNeu-TTS-1000h. 2x faster, ultra-low latency.
## Technical Details
| Spec | Value |
|------|-------|
| Training data | VieNeu-TTS-1000h (443,641 samples) |
| Audio codec | NeuCodec |
| Context window | 2,048 tokens |
| Output sample rate | 24 kHz |
| Watermark | Enabled by default |
## HuggingFace Links
- [VieNeu-TTS (0.5B)](https://huggingface.co/pnnbao-ump/VieNeu-TTS)
- [VieNeu-TTS-0.3B](https://huggingface.co/pnnbao-ump/VieNeu-TTS-0.3B)
- [Training Dataset](https://huggingface.co/datasets/pnnbao-ump/VieNeu-TTS-1000h)