# VieNeu documentation — full text > Vietnamese Text-to-Speech with Instant Voice Cloning Every documentation page below, in sidebar order, each headed by the URL it is served from. The API reference is not included here — it is generated separately at https://docs.vieneu.io/api-reference.md, from https://docs.vieneu.io/openapi.json. Index: https://docs.vieneu.io/llms.txt --- # Introduction Source: https://docs.vieneu.io/docs/ VieNeu turns Vietnamese text into natural speech. :::tip Using Claude, ChatGPT or Cursor? No code needed Add `https://api.vieneu.io/mcp` to your AI assistant, sign in with your VieNeu account, and ask it to read in a Vietnamese voice — the audio plays right in the chat. Cursor and VS Code install it in one click. **[Set it up in a minute →](./integrations/mcp/index.md)** ::: To build VieNeu into your own code, there are **two ways to use it**, and they are different products — start by picking one. ## Which one do you want? ### ☁️ Cloud API — `api.vieneu.io` Send text over HTTPS, get audio back. Nothing to install, no GPU, no model download. Billed per character. This is what you want if you are **adding a Vietnamese voice to a product**: an app, a website, a voice agent, a dubbing pipeline. There is an OpenAI-compatible endpoint, so if your code already calls OpenAI's `/v1/audio/speech`, pointing it at VieNeu is a base URL and an API key. - **[Quickstart](./cloud-api/quickstart.md)** — a key, a voice, one call, audio out - **[Cloud API overview](./cloud-api/overview.md)** — auth, billing, every endpoint - **[OpenAI-compatible endpoint](./cloud-api/openai-compatible.md)** — the fastest way in ### 💻 On-device SDK — the Python package Runs the model on your own machine. No network call per request, no per-character cost, and the text never leaves your hardware — in exchange you supply the hardware and the setup. Documented in the **SDK** and **Getting Started** sections of this site, and summarized below. :::info They are separate products The Cloud API and the SDK share a name and a voice catalogue. They do **not** share an interface: different install, different authentication, different request shapes, different billing. Code written against one does not run against the other, and neither section's documentation applies to the other. Pick the one you are actually using and stay in it. ::: --- ## The on-device SDK **VieNeu-TTS** is an advanced on-device Vietnamese Text-to-Speech (TTS) system with **instant voice cloning**. Give it text, it speaks it back in natural Vietnamese — fully offline, no cloud API needed. ## Key Features - **Instant Voice Cloning** — Clone any voice with just 3-5 seconds of reference audio - **Code-switching** — Seamless transitions between Vietnamese and English - **Real-time Streaming** — Start audio playback before the entire sentence is finished - **Multiple Backends** — PyTorch (GPU), GGUF quantized (CPU), LMDeploy (fast GPU), Remote API - **Production Ready** — 24 kHz waveform generation, audio watermarking ## How It Works VieNeu-TTS uses a **causal language model** to generate speech. The core pipeline: ``` Text → Normalize → Phonemize (eSpeak NG) → LLM generates speech tokens → Codec decodes to audio ``` 1. **Text normalization** — Converts numbers, abbreviations, punctuation to spoken form 2. **Phonemization** — eSpeak NG converts text to pronunciation symbols 3. **Token generation** — A transformer LLM predicts discrete speech tokens 4. **Audio decoding** — NeuCodec converts tokens into a 24kHz waveform ## Models | Model | Format | Quality | Speed | |-------|--------|---------|-------| | VieNeu-TTS (0.5B) | PyTorch | Best | Very Fast (GPU) | | VieNeu-TTS-0.3B | PyTorch | Great | Ultra Fast (2x) | | GGUF Q8 variants | GGUF | Great | Fast (CPU) | | GGUF Q4 variants | GGUF | Good | Very Fast (CPU) | All models are hosted on [HuggingFace](https://huggingface.co/pnnbao-ump) and auto-downloaded on first use. ## Quick Start ```bash git clone https://github.com/pnnbao97/VieNeu-TTS.git cd VieNeu-TTS uv sync uv run vieneu-web ``` Open `http://127.0.0.1:7860` and start generating speech. --- # MCP server — VieNeu in Claude, ChatGPT and more Source: https://docs.vieneu.io/docs/integrations/mcp/ VieNeu runs a hosted [Model Context Protocol](https://modelcontextprotocol.io) (MCP) server. Add one URL to an AI assistant, sign in with your VieNeu account, and ask for Vietnamese speech in plain language — the assistant finds a voice, synthesizes the text, and plays it back or hands you a download link. ``` https://api.vieneu.io/mcp ``` Nothing to install. The assistant signs in through your browser (OAuth 2.1) the first time you use it; tools and apps that cannot sign in can send an [API key](#using-an-api-key-instead-of-signing-in) instead. **Before you start:** you need a VieNeu account with an **active token plan**. Connecting is refused without one, because every synthesis would fail. ## Pick your app | App | How you connect | Inline player | Guide | |---|---|---|---| | **Claude** (web, desktop, mobile) | Add a custom connector, sign in | Yes | [Claude](./claude-ai.md) | | **ChatGPT** | Developer mode → add an app, sign in | Yes | [ChatGPT](./chatgpt.md) | | **Claude Code** | One command, then `/mcp` | — | [Claude Code](./claude-code.md) | | **Cursor** | One click or `mcp.json`; sign in or API key | Yes | [Cursor](./cursor.md) | | **VS Code** (GitHub Copilot) | One click or `mcp.json`; sign in or API key | Yes (experimental) | [VS Code](./vscode.md) | | **OpenAI API / Agents SDK** | `mcp` tool in your code | — | [OpenAI API](./openai-api.md) | | Windsurf (Devin Desktop), Gemini CLI, Codex CLI, Zed, Goose, LM Studio | Config file | Goose only | [Other apps](./other-clients.md) | Something not working? See [Troubleshooting](./troubleshooting.md). ## What you say to it > "What can VieNeu do?" The assistant calls `list_capabilities` and answers with what works here, one example request for each, and which ones spend tokens. In apps with the [inline player](#the-inline-player) it is a card: click an example and it is sent to the chat. > "Find me a northern female voice for news." The assistant calls `list_voices`. Free. > "Read this paragraph with Thu Trang." The assistant calls `text_to_speech`. In apps with the inline player, a small player appears right in the chat — press play. There is always a download link too (valid 24 hours; MP3 for texts up to 1,500 characters). The audio is saved in your **Library** on vieneu.io. This spends tokens from your plan, exactly like the same request made through the API. > "Add a laugh after the first sentence." The assistant calls `list_emotion_tags` and writes a cue tag such as `[cười]` into the text before synthesizing. > "Make an audiobook of this story — a northern woman narrating, a voice for each character." The assistant scripts your text, chooses the voices, tells you the cost, and after you agree VieNeu makes one MP3 per chapter in the background. See [Audiobooks](./audiobooks.md). ## Tools | Tool | What it does | Cost | |------|--------------|------| | `list_capabilities` | What this connection can do, with an example request for each | Free | | `list_voices` | Search the voices your account can use, including your own cloned voices | Free | | `text_to_speech` | Synthesize Vietnamese text; returns a download link and the duration | Tokens | | `get_speech_status` | Check a long job that was still processing; returns the link when done | Free | | `list_emotion_tags` | Reading styles, and cue tags such as `[cười]` to write into the text | Free | | `get_token_balance` | Tokens you can spend right now, your plan's caps and when they reset, and roughly how many characters that reads | Free | Six more tools — `create_audiobook`, `add_audiobook_chapter`, `start_audiobook`, `get_audiobook`, `list_audiobooks`, `cancel_audiobook` — make audiobooks; they are described in [Audiobooks](./audiobooks.md#tools). ### `list_voices` | Argument | Type | Notes | |---|---|---| | `search` | string, optional | A few words describing the voice (`nữ miền Bắc trầm ấm`, `male south`), or its name (`Thu Trang`). How the words are read is below. | | `limit` | 1–100, optional | Default 25. The reply says how many more matched. | Each line reads `id (gender, region) "Name" - description`. The **id** is what `text_to_speech` needs, exactly as shown — ids look like names (`Thu Trang`) and cannot be guessed. How `search` reads the words: - **Gender and region** come from the voice's fields, never from its description: `nữ`, `nam`, `miền Bắc`, `miền Nam`, `miền Trung`, or `female`, `male`, `north`, `south`, `central`. `nam` on its own means male; write `miền Nam` for the South. - **Other words** match whole words of the voice's id, name or description, ignoring case. A word written with diacritics must match them (`già`, old, does not find a voice named Gia); written without, it matches either way, and from 4 letters also as the start of a word. - **Words that describe no voice**, such as `giọng` or `đọc`, are ignored. - **A voice's name** (`Thùy An`, `giọng Thu Trang`) puts that voice first. - **When no voice has every word**, the closest come back instead, those with the gender and region asked for first, and each line ends with the words that voice lacks: `[thiếu: trầm, ấm]`. ### `text_to_speech` | Argument | Type | Notes | |---|---|---| | `text` | string | Vietnamese text, up to 50,000 characters. Numbers, dates and English acronyms are handled. | | `voice` | string, optional | A voice id from `list_voices`. Omitted: the default voice. | | `speed` | 0.5–2.0, optional | 1.0 is natural pace. | | `style` | string, optional | A reading style from `list_emotion_tags`. | Up to 1,500 characters the tool makes one [`POST /v1/audio/speech`](../../cloud-api/openai-compatible.md) call for **MP3** and returns its link a moment later. Longer text becomes one [`POST /v1/tts`](../../cloud-api/overview.md) job (**WAV**) that the tool waits on for about 50 seconds; a very long text may still be processing then — the reply gives a job id and tells the assistant to call `get_speech_status` later instead of synthesizing again (which would be charged again). Either way the reply carries the link, its format, the duration and the voice — and, once the audio is ready, how many tokens you have left. ### `get_speech_status` | Argument | Type | Notes | |---|---|---| | `job_id` | string | The id `text_to_speech` returned. | ### `get_token_balance` No arguments. Ask *"còn bao nhiêu token?"* and the assistant answers with: - **Usable now** — the most one synthesis can spend at this moment: the plan's remaining tokens, or less when a daily or weekly cap is lower. - The plan's remaining / total, each cap and when it resets (Vietnam time), and when the plan expires. - The total across all your active plans, if you have more than one — the next one takes over when this one runs out. - About how many characters that reads with the default voice engine. At zero it links to the pricing page. The same numbers are on [`GET /v1/balance`](../../cloud-api/overview.md#seeing-what-you-were-billed-for) for scripts. ## The inline player In apps that support [MCP Apps](https://modelcontextprotocol.io/docs/extensions/apps) — Claude (web, desktop and mobile), ChatGPT, Cursor, VS Code (experimental) and Goose — `text_to_speech` and `get_speech_status` results show a player in the conversation: play/pause, seek, duration, and a download button. While a long job is still running it shows a small green light and keeps checking `get_speech_status` by itself (free), switching to the player when the audio is ready. Audiobooks get their own player under `start_audiobook` and `get_audiobook`: the chapters with their progress, played one after another, refreshing itself while the book is being made. Claude asks once before showing it — choose **Allow** (or **Always allow**). Apps without MCP Apps show the text reply with the link instead; nothing else changes. ## Ready-made prompts {#prompts} The server also offers five prompts — ready-made requests with a form for their details. Apps that show MCP prompts list them in a menu: Claude Code as `/mcp__vieneu__make_audiobook`, VS Code as `/mcp.vieneu.make_audiobook` (the middle part is the name you gave the server). ChatGPT does not show prompts; ask *"VieNeu làm được gì?"* there instead. | Prompt | What it asks for | Details | |---|---|---| | `what_can_vieneu_do` | The list of what VieNeu can do | — | | `read_text` | Read a text aloud | `text`, `voice` (optional, e.g. "nữ miền Bắc") | | `make_audiobook` | An audiobook from your text, priced before it starts | `text` (optional — paste it after), `narrator` (optional) | | `find_voice` | Voices matching a description | `description` | | `check_balance` | Tokens left and the plan's limits | — | ## Using an API key instead of signing in Scripts, the OpenAI API, and apps that cannot sign in can send one of your [API keys](../../cloud-api/overview.md#authentication) on every request, in either header: ``` X-API-Key: vn_sk_... Authorization: Bearer vn_sk_... ``` Create a key on the **Developer** page of vieneu.io. A config file holding a key is a password on disk — prefer signing in where the app supports it, and never commit a key to a repository. Each app's guide shows where the header goes. ## Billing A tool call costs exactly what the equivalent API call costs: one synthesis per `text_to_speech` call, charged per character (minimum 50) at the engine's rate, from your plan's tokens. A failed synthesis is refunded automatically. An audiobook is billed the same way, chapter by chapter as each one is made, after you agree to `start_audiobook`. Every other tool is free. MCP usage appears under **Developer → Usage** like any API traffic, attributed to the connection (for example "Claude (MCP)"). ## See and remove connections **Developer → Dùng VieNeu trong Claude, ChatGPT** on vieneu.io shows the URL to copy and lists every app you have connected, with the date and the last use. **Gỡ kết nối** (Disconnect) signs that app out immediately; it has to ask for permission again to come back. ## Errors The tools translate API errors into sentences the assistant can act on. | What you see | Cause | Fix | |---|---|---| | Asked to sign in again | The connection was removed on vieneu.io, or its sign-in expired | Reconnect from the app | | "Tài khoản không đủ quyền hoặc hết token…" | The plan is out of tokens, expired, or does not include the engine (HTTP 403) | Top up or renew on vieneu.io — retrying does not help | | "Đang gọi quá nhanh…" | Rate limited (HTTP 429) | Wait the number of seconds quoted | | "Máy chủ tạo giọng đang bận" | No synthesis worker free (HTTP 503) | Retry in a few seconds | | "Voice … is not available" | The voice id is not in your catalogue | Pick an id from `list_voices` | More in [Troubleshooting](./troubleshooting.md). ## What is deliberately not exposed Voice cloning, dubbing and SRT dubbing are not tools here. They take file uploads and spend noticeably more per call, and an assistant invoking one by accident is a bad first experience. Use the web app or the [Cloud API](../../cloud-api/overview.md) for those. ## For client developers - Transport: Streamable HTTP, stateless. Protocol revision 2026-07-28, with 2025-era clients served too. - Authorization: OAuth 2.1 per the MCP authorization spec. An unauthenticated request gets `401` with `WWW-Authenticate: Bearer resource_metadata="https://api.vieneu.io/.well-known/oauth-protected-resource"`. - Metadata: [`/.well-known/oauth-protected-resource`](https://api.vieneu.io/.well-known/oauth-protected-resource) (RFC 9728) and [`/.well-known/oauth-authorization-server`](https://api.vieneu.io/.well-known/oauth-authorization-server) (RFC 8414). - Clients: Client ID Metadata Documents and Dynamic Client Registration (`/oauth/register`). PKCE `S256` is always required. A metadata document must allow `none` (preferred or among `token_endpoint_auth_methods_supported`). Registration accepts `none`, `client_secret_post` and `client_secret_basic` and always returns a `client_secret`; it is required only for the two secret methods. Loopback redirects (`http://127.0.0.1` / `localhost`) match on any port. - Scope: `tts` (plus `offline_access` for a refresh token). Access tokens last one hour; refresh tokens rotate on every use. - MCP Apps: `ui://vieneu/player.html`, `ui://vieneu/audiobook.html` and `ui://vieneu/capabilities.html` (`text/html;profile=mcp-app`); media allowed from the storage origin only. The capabilities card sends an example to the chat with `ui/message` when the host offers it. - Prompts: `what_can_vieneu_do`, `read_text`, `make_audiobook`, `find_voice`, `check_balance`. --- # Make an audiobook with Claude or ChatGPT Source: https://docs.vieneu.io/docs/integrations/mcp/audiobooks Give the assistant a story, a chapter or a whole book — paste it or attach the file — and ask for an audiobook: > "Làm sách nói từ truyện này. Giọng dẫn chuyện nữ miền Bắc, mỗi nhân vật một giọng riêng." The assistant turns your text into a script — the narration, and every line of dialogue with the character who says it — picks a narrator and a voice for each character, and tells you what it will cost. Once you agree, VieNeu reads the book chapter by chapter and masters each chapter into an MP3 ready to publish. This works in every app on the [MCP server](./index.md) — Claude and ChatGPT show a player for the chapters right in the chat. ## How it goes 1. **The script.** The assistant creates the book and adds the chapters in order. It copies your text word for word — its instructions forbid summarizing or rewriting — and leaves out tables of contents, page numbers and copyright pages. Free. 2. **The cost.** Before anything is spent it tells you the estimate in tokens: every character is billed like any synthesis, with the default engine's rate (`get_token_balance` tells you how far your plan goes). 3. **Start.** When you say yes it starts the book. Each chapter is billed when it starts rendering; a chapter that fails is refunded automatically and made again on its own (see [failures](#stopping-resuming-failures)). 4. **The wait.** Audiobooks are made while VieNeu's servers are not busy with people waiting for audio — expect hours rather than minutes, often overnight. Books waiting at the same time take turns, a chapter each, so a short book is never stuck behind a long one. You do not need to keep the chat open: when the book is done, a notification appears under the bell on vieneu.io. 5. **Listening.** Ask *"Sách nói xong chưa?"* The assistant shows the progress and, for each finished chapter, a download link (valid 24 hours — ask again for fresh ones). In Claude and ChatGPT a player lists the chapters and plays them one after another. Every chapter is also saved in your **Library** on vieneu.io. ## Voices - **The narrator** reads the narration and the chapter titles. - **Characters** — up to 30 — each get their own voice. Minor characters can stay with the narrator: the assistant writes their lines as narration. - Describe what you want ("một bà cụ miền Nam", "cậu bé tám tuổi") and the assistant searches the catalogue with `list_voices`. Your own cloned voices cannot read audiobooks yet. ## How a chapter sounds | What | How | |---|---| | Between paragraphs and lines | 0.6 s of silence | | After the chapter title | 2.5 s | | At a scene change | 2 s | | Start and end of the file | 0.5 s of silence before, 3.5 s after | | Loudness | −19 LUFS integrated, peaks at most −3 dBTP | | File | MP3, 128 kbps, mono, 48 kHz, tagged with title, book, author and track number | Stores differ in the files they accept — Audible, for one, asks for 192 kbps at 44.1 kHz — so check your store's requirements before uploading. ## Tools | Tool | What it does | Cost | |------|--------------|------| | `create_audiobook` | Create the book: title, author, narrator, characters' voices | Free | | `add_audiobook_chapter` | Add a chapter as a script in reading order, or continue a long one | Free¹ | | `start_audiobook` | Start the book — also resumes a stopped book and redoes failed chapters | Tokens | | `get_audiobook` | Progress, and the download link of each finished chapter | Free | | `list_audiobooks` | Your books, newest first — finds a book from an earlier chat | Free | | `cancel_audiobook` | Stop: chapters not started are never billed; one already being made finishes | Free | ¹ A chapter added to a book that is already being made joins it, and is billed when it is made. A chapter's script is a list of segments: narration (no speaker), or one character's spoken words. A line such as *"— Đi thôi! — Lan nói."* becomes Lan saying *"Đi thôi!"* followed by the narrator reading *"Lan nói."* ## Limits | Limit | Value | |---|---| | Chapters per book | 200 | | Characters per chapter | 60,000 (about 80 minutes), less on a plan with a daily cap¹ | | Characters per book | 1,500,000 | | Characters with their own voice | 30 | ¹ A chapter is billed in one go, so it has to fit the plan's daily (or weekly) token cap — on a trial of 50,000 tokens a day that is about 12,800 characters. The book reports its ceiling (`chapter_chars_max`), and a chapter over it is refused when added, with the size that fits, rather than waiting for tokens that never come. The assistant splits such a chapter into several. An assistant writes every word of the book into its tool calls, a few thousand characters at a time, so long chapters arrive in several pieces — that is normal. Each piece says how many segments the chapter already has (`after_segment`): a piece sent twice is refused instead of being added twice, a piece that went missing is noticed before the next one lands, and each answer shows the words the chapter now ends with. Before starting, the assistant checks every chapter against your text. A whole novel is a long conversation: it is fine to carry on in a new chat, where *"tiếp tục sách nói …"* finds the book again. ## Stopping, resuming, failures - **"Dừng sách nói"** stops the book. Chapters not started are not billed; the one being made finishes, because it is already paid for. - **"Làm tiếp"** resumes it, with the cost of what is left. - A chapter whose rendering fails — a server restarting mid-chapter, a voice server down — is refunded and made again on its own a few minutes later, continuing from the parts already made, twice. Only a third failure counts: the book then ends as *partly done*, and starting it again redoes only the failed chapters. - A chapter whose job stops moving for 45 minutes is treated the same way, so one stuck chapter never holds up the books behind it. - If the finishing step (the MP3) fails, it is tried again with growing waits for about an hour; only then is the chapter handed over as WAV instead. ## Your text You are reading your text to VieNeu, so make sure you may: your own writing, a public-domain work, or one you hold the audio rights to. ## From your own code The same books are in the Cloud API: `POST /v1/audiobooks`, then `POST /v1/audiobooks/{bookId}/chapters` and `POST /v1/audiobooks/{bookId}/start`; follow progress with `GET /v1/audiobooks/{bookId}`. See [Every operation](../../cloud-api/overview.md#every-operation) and the [API reference](/api-reference). --- # Use VieNeu in Claude Source: https://docs.vieneu.io/docs/integrations/mcp/claude-ai Works in Claude on the **web** (claude.ai), **Claude Desktop** and the **Claude mobile** apps — one connector, added once, available everywhere you sign in to Claude. The audio plays in a small [player](./index.md#the-inline-player) right under the reply. **You need:** a Claude account (Free plans can add **one** custom connector; Pro and Max can add more) and a VieNeu account with an active token plan. ## Add the connector (Free, Pro, Max) 1. In Claude, open **Customize → Connectors** ([claude.ai/customize/connectors](https://claude.ai/customize/connectors)). 2. Click **Add custom connector** (on some screens: **+** first). 3. Fill in: - **Name:** `VieNeu` - **URL:** `https://api.vieneu.io/mcp` Leave any OAuth client fields empty — Claude identifies itself to VieNeu on its own. 4. Click **Add**, then **Connect**. 5. A VieNeu page opens. Sign in if asked, check that the **account** shown is the one you want, and press **Cho phép** (Allow). You land back in Claude. Add it on the web or in Claude Desktop; the mobile apps then use the same connector. ## Add the connector (Team, Enterprise) Only an organization **Owner** can add a custom connector. 1. The Owner opens **Organization settings → Connectors**, chooses **Add → Custom** (then **Web** if asked), enters `https://api.vieneu.io/mcp` and clicks **Add**. 2. Each member then opens **Customize → Connectors**, finds **VieNeu** (marked *Custom*), clicks **Connect**, and approves their **own** VieNeu account on the VieNeu page. Each member's usage is billed to their own VieNeu plan. ## Use it in a chat 1. In the chat box, click **+** → **Connectors** and make sure **VieNeu** is switched on for this conversation. 2. Ask in plain language, for example: - "Tìm giọng nữ miền Bắc rồi đọc câu: Xin chào, đây là VieNeu." - "Đọc đoạn này bằng giọng Thu Trang, tốc độ 1.1." 3. The first time, Claude asks permission to use a VieNeu tool — choose **Allow once** or **Always allow**. It also asks once before showing the player — choose **Allow** (or **Always allow**). Each `text_to_speech` call spends tokens. Claude may offer a couple of voices to compare; each one it reads is a separate charge. ## Using an API key instead Claude's connector dialog can send a fixed request header only for a limited set of organizations (a beta feature, Owners only). If your organization has a **Request headers** section when adding the connector, choose **No sign-in** and add the header `x-api-key` with your key, or `authorization` with the value `Bearer vn_sk_...`. The key is shared by everyone in the organization — use sign-in unless you specifically need this. ## Remove it - From Claude: **Customize → Connectors → VieNeu → Remove** (or disconnect). - From VieNeu: **Developer → Dùng VieNeu trong Claude, ChatGPT → Gỡ kết nối**. This signs Claude out immediately, even on other devices. ## Problems | Symptom | Fix | |---|---| | "Couldn't reach the MCP server" when adding | Check the URL is exactly `https://api.vieneu.io/mcp` (no trailing slash, no spaces) | | The VieNeu page says the request expired | Go back to Claude and click **Connect** again — a request is valid for 30 minutes | | "Tài khoản chưa có gói token còn hạn" on the VieNeu page | Buy or renew a plan on vieneu.io, then connect again | | No player, only a link | Claude asks once before showing it — choose **Allow**. The link always works | | Claude says it has no VieNeu tools | Turn the connector on for the conversation (**+ → Connectors**) | More in [Troubleshooting](./troubleshooting.md). --- # Use VieNeu in ChatGPT Source: https://docs.vieneu.io/docs/integrations/mcp/chatgpt ChatGPT connects to custom MCP servers through **developer mode**. The audio plays in the [inline player](./index.md#the-inline-player) under the reply. **You need:** ChatGPT **Plus, Pro, Business, Enterprise or Edu**, on **chatgpt.com** (the web). The Free plan cannot add custom apps. On Business and Enterprise, a workspace admin may have to allow developer mode or custom apps first. And a VieNeu account with an active token plan. :::note ChatGPT's menus for custom apps have been renamed several times. The steps below match OpenAI's guide as of October 2026; if a label differs, look for "Developer mode" and "add app / connector" in the same area. ::: ## Add the app 1. Open **Settings → Security and login** and turn on **Developer mode**. 2. Open **Plugins** (apps) and click **+** to create one. 3. Fill in: - **Name:** `VieNeu` - **Description:** `Vietnamese text-to-speech` - **Server URL / Connection:** `https://api.vieneu.io/mcp` - **Authentication:** **OAuth** 4. Create it. ChatGPT registers itself with VieNeu automatically and opens the VieNeu page: sign in if asked, check the account and press **Cho phép** (Allow). 5. ChatGPT lists the five VieNeu tools. The app appears under **Drafts**. ChatGPT does not support API keys for custom apps — use sign-in. ## Use it in a chat 1. In a new chat, open **+ → Developer mode** and select **VieNeu**. 2. Ask, for example: *"Tìm giọng nam miền Nam đọc tin tức, rồi đọc đoạn sau…"* 3. ChatGPT asks you to confirm before a tool that makes something (like `text_to_speech`) runs — confirm it. Each synthesis spends tokens. ## Remove it - In ChatGPT: remove the app from **Plugins** (or turn developer mode off). - On VieNeu: **Developer → Dùng VieNeu trong Claude, ChatGPT → Gỡ kết nối**. ## Problems | Symptom | Fix | |---|---| | No "Developer mode" option | Your plan or workspace does not allow it — see *You need* above | | The VieNeu page warns the app "calls itself ChatGPT" but goes elsewhere | Do **not** allow it: the request did not come from ChatGPT. Real ChatGPT returns to `chatgpt.com` | | Tools are missing after an update | Open the app under **Drafts** and click **Refresh** | More in [Troubleshooting](./troubleshooting.md). --- # Use VieNeu in Claude Code Source: https://docs.vieneu.io/docs/integrations/mcp/claude-code ## Add the server ```bash claude mcp add --transport http vieneu https://api.vieneu.io/mcp ``` Then, inside Claude Code, run `/mcp`, pick **vieneu** and choose **Authenticate**. Your browser opens the VieNeu consent page. It shows a yellow note that the app runs on your computer (the callback is `127.0.0.1` or `localhost`) — that is expected for Claude Code. Press **Cho phép** (Allow). From a shell you can also start the sign-in with: ```bash claude mcp login vieneu ``` ### Where it is saved (`--scope`) | Scope | Stored in | Use when | |---|---|---| | `local` (default) | `~/.claude.json`, this project only | Just you, just here | | `user` | `~/.claude.json`, every project | You want VieNeu everywhere | | `project` | `.mcp.json` in the repository | The whole team should get it (each person still signs in) | ```bash claude mcp add --transport http --scope user vieneu https://api.vieneu.io/mcp ``` ## Using an API key instead For CI or machines where a browser sign-in is not possible: ```bash claude mcp add --transport http vieneu https://api.vieneu.io/mcp \ --header "X-API-Key: vn_sk_..." ``` In a shared `.mcp.json`, reference an environment variable instead of the key: ```json { "mcpServers": { "vieneu": { "type": "http", "url": "https://api.vieneu.io/mcp", "headers": { "X-API-Key": "${VIENEU_API_KEY}" } } } } ``` ## Use it Ask in the session, for example: *"Đọc README tiếng Việt này bằng giọng nữ miền Nam và cho mình link tải."* Claude Code shows the reply and the download link; it has no inline player. ## Remove it ```bash claude mcp remove vieneu ``` and, if you signed in, **Developer → Gỡ kết nối** on vieneu.io. --- # Use VieNeu in Cursor Source: https://docs.vieneu.io/docs/integrations/mcp/cursor ## Add the server **One click:** **Add VieNeu to Cursor**. Cursor opens with the server filled in; confirm **Install**, then sign in (below). Or by hand: create or edit `mcp.json` — **`.cursor/mcp.json`** in a project, or **`~/.cursor/mcp.json`** for every project: ```json { "mcpServers": { "vieneu": { "url": "https://api.vieneu.io/mcp" } } } ``` Save it. Cursor lists **vieneu** under its MCP settings and asks you to sign in: your browser opens the VieNeu page — press **Cho phép** (Allow). The page notes the app runs on your computer; that is expected for Cursor. ## Using an API key instead ```json { "mcpServers": { "vieneu": { "url": "https://api.vieneu.io/mcp", "headers": { "X-API-Key": "${env:VIENEU_API_KEY}" } } } } ``` `${env:VIENEU_API_KEY}` reads the key from your environment, so the file can be committed without the secret. ## Use it In Agent chat: *"Đọc đoạn mô tả sản phẩm này bằng giọng nữ miền Nam, trả link MP3."* Cursor asks before running a tool (you can allow it). Results render in the [inline player](./index.md#the-inline-player) where Cursor supports MCP Apps; the link is always in the reply. ## Teams On Cursor Business/Enterprise, admins can restrict which MCP servers members may use (**Team Settings → MCP Configuration**). Ask your admin to allow `https://api.vieneu.io/mcp`. --- # Use VieNeu in VS Code with GitHub Copilot Source: https://docs.vieneu.io/docs/integrations/mcp/vscode ## Add the server **One click:** **Add VieNeu to VS Code**. VS Code opens the server's install page; choose **Install**, then sign in when it starts. Or run **MCP: Add Server** from the Command Palette and choose **HTTP**, or create **`.vscode/mcp.json`** in your workspace (or run **MCP: Open User Configuration** for every workspace): ```json { "servers": { "vieneu": { "type": "http", "url": "https://api.vieneu.io/mcp" } } } ``` The first time the server starts, VS Code asks whether you trust it, then opens your browser for sign-in: press **Cho phép** (Allow) on the VieNeu page. ## Using an API key instead VS Code can prompt for the key once and store it securely: ```json { "inputs": [ { "type": "promptString", "id": "vieneu-key", "description": "VieNeu API key", "password": true } ], "servers": { "vieneu": { "type": "http", "url": "https://api.vieneu.io/mcp", "headers": { "X-API-Key": "${input:vieneu-key}" } } } } ``` ## Use it Open Copilot Chat in **Agent** mode, make sure the VieNeu tools are enabled in the tools picker, and ask: *"Đọc changelog này bằng giọng nam miền Bắc."* **Inline player:** VS Code shows MCP Apps behind an experimental setting. Turn on `chat.mcp.apps.enabled` to get the [player](./index.md#the-inline-player) in chat; otherwise you get the link. ## Copilot Business and Enterprise Organizations must enable the **"MCP servers in Copilot"** policy (off by default) before members can use any MCP server. Copilot Free, Pro and Pro+ are not affected. --- # Use VieNeu from the OpenAI API Source: https://docs.vieneu.io/docs/integrations/mcp/openai-api If you build your own app on OpenAI models, you can hand the model the VieNeu tools directly: the model decides when to search voices and synthesize, and OpenAI calls `https://api.vieneu.io/mcp` for you. **You need:** an OpenAI API key and a VieNeu [API key](../../cloud-api/overview.md#authentication) (`vn_sk_...`, from the **Developer** page). Keep both in environment variables: ```bash export OPENAI_API_KEY=sk-... export VIENEU_API_KEY=vn_sk_... ``` :::tip Just want audio from code? If your code already knows the text and the voice, you do not need MCP or an LLM at all — call the [OpenAI-compatible endpoint](../../cloud-api/openai-compatible.md) directly with the OpenAI SDK. MCP is for letting the *model* decide. ::: ## Responses API The Responses API has a built-in remote MCP tool. Pass your VieNeu key in `authorization`; OpenAI sends it to VieNeu as a Bearer token, which VieNeu accepts. ```python import os from openai import OpenAI client = OpenAI() resp = client.responses.create( model="gpt-5", # any current model that supports tools tools=[{ "type": "mcp", "server_label": "vieneu", "server_url": "https://api.vieneu.io/mcp", "authorization": os.environ["VIENEU_API_KEY"], "allowed_tools": ["list_voices", "text_to_speech", "get_speech_status", "get_token_balance"], "require_approval": "never", }], input="Tìm một giọng nữ miền Bắc rồi đọc câu: Xin chào, đây là VieNeu. Trả về link tải.", ) print(resp.output_text) ``` ```js import OpenAI from "openai"; const client = new OpenAI(); const resp = await client.responses.create({ model: "gpt-5", tools: [{ type: "mcp", server_label: "vieneu", server_url: "https://api.vieneu.io/mcp", authorization: process.env.VIENEU_API_KEY, allowed_tools: ["list_voices", "text_to_speech", "get_speech_status", "get_token_balance"], require_approval: "never", }], input: "Tìm một giọng nữ miền Bắc rồi đọc câu: Xin chào, đây là VieNeu. Trả về link tải.", }); console.log(resp.output_text); ``` - **`require_approval`:** the default asks for approval before every tool call (the response then contains `mcp_approval_request` items you must answer). `"never"` lets the model spend your VieNeu tokens without asking — keep `allowed_tools` tight, or require approval for `text_to_speech` only: `"require_approval": {"always": {"tool_names": ["text_to_speech"]}}`. - **The key is not stored by OpenAI**, so send it with every request. - The audio link is in the tool result and usually in the model's answer; the `mcp_call` items in `resp.output` hold the raw tool results. ## OpenAI Agents SDK (Python) Let OpenAI call VieNeu for you (hosted tool): ```python import os from agents import Agent, HostedMCPTool, Runner agent = Agent( name="Narrator", instructions="You narrate Vietnamese text with VieNeu and return the audio link.", tools=[HostedMCPTool(tool_config={ "type": "mcp", "server_label": "vieneu", "server_url": "https://api.vieneu.io/mcp", "authorization": os.environ["VIENEU_API_KEY"], "require_approval": "never", })], ) result = Runner.run_sync(agent, "Đọc câu 'Chào buổi sáng' bằng giọng Thu Trang.") print(result.final_output) ``` Or connect from your own process (your code calls VieNeu, with any header you choose): ```python import asyncio, os from agents import Agent, Runner from agents.mcp import MCPServerStreamableHttp async def main(): async with MCPServerStreamableHttp( name="vieneu", params={ "url": "https://api.vieneu.io/mcp", "headers": {"X-API-Key": os.environ["VIENEU_API_KEY"]}, "timeout": 60, }, cache_tools_list=True, ) as vieneu: agent = Agent(name="Narrator", mcp_servers=[vieneu]) result = await Runner.run(agent, "Đọc câu 'Chào buổi sáng' bằng giọng Thu Trang.") print(result.final_output) asyncio.run(main()) ``` Set `timeout` generously: `text_to_speech` waits for the audio (a few seconds for short text, up to about 50 seconds for long text). ## Costs Two bills: OpenAI charges for the model's tokens, VieNeu charges your plan for each synthesis (per character, minimum 50). `list_voices`, `list_emotion_tags`, `get_speech_status` and `get_token_balance` are free on the VieNeu side. --- # Use VieNeu in other MCP apps Source: https://docs.vieneu.io/docs/integrations/mcp/other-clients Any app that supports **remote MCP servers over Streamable HTTP** works with `https://api.vieneu.io/mcp`. Apps that support MCP sign-in open the VieNeu page on their own; the rest can send an [API key](./index.md#using-an-api-key-instead-of-signing-in). Replace `vn_sk_...` with your key, or better, keep it in an environment variable `VIENEU_API_KEY` as shown. ## Windsurf (Devin Desktop) Windsurf is now called **Devin Desktop**. Edit its `mcp_config.json` (open it from the MCP settings in the app — the file's location changed with the rename). Note the key is `serverUrl`, not `url`: ```json { "mcpServers": { "vieneu": { "serverUrl": "https://api.vieneu.io/mcp", "headers": { "X-API-Key": "${env:VIENEU_API_KEY}" } } } } ``` Leave out `headers` to sign in instead. ## Gemini CLI ```bash gemini mcp add --transport http vieneu https://api.vieneu.io/mcp ``` Then sign in inside the CLI with `/mcp auth vieneu`. Or edit `~/.gemini/settings.json` (`httpUrl` means Streamable HTTP): ```json { "mcpServers": { "vieneu": { "httpUrl": "https://api.vieneu.io/mcp", "headers": { "X-API-Key": "vn_sk_..." } } } } ``` ## OpenAI Codex CLI ```bash codex mcp add vieneu --url https://api.vieneu.io/mcp codex mcp login vieneu ``` Or with an API key, in `~/.codex/config.toml`: ```toml [mcp_servers.vieneu] url = "https://api.vieneu.io/mcp" bearer_token_env_var = "VIENEU_API_KEY" # sends Authorization: Bearer ``` ## Zed In Zed's `settings.json`: ```json { "context_servers": { "vieneu": { "url": "https://api.vieneu.io/mcp", "headers": { "Authorization": "Bearer vn_sk_..." } } } } ``` Without an `Authorization` header Zed starts the sign-in flow instead. ## Goose `goose configure` → **Add Extension** → **Remote Extension (Streamable HTTP)** → URL `https://api.vieneu.io/mcp` (in Goose Desktop: **Extensions → Add custom extension**). Goose signs in on its own, or you can add an `X-API-Key` header in the wizard. Goose shows the [inline player](./index.md#the-inline-player). ## LM Studio LM Studio uses a token header (no sign-in). In its `mcp.json`: ```json { "mcpServers": { "vieneu": { "url": "https://api.vieneu.io/mcp", "headers": { "Authorization": "Bearer vn_sk_..." } } } } ``` ## An app not listed here Look for "remote MCP server", "HTTP" or "Streamable HTTP" in its settings and use the URL. If it only supports local (stdio) servers, bridge it with [`mcp-remote`](https://www.npmjs.com/package/mcp-remote): ```json { "mcpServers": { "vieneu": { "command": "npx", "args": ["-y", "mcp-remote", "https://api.vieneu.io/mcp"] } } } ``` `mcp-remote` opens the VieNeu sign-in page in your browser the first time. --- # MCP troubleshooting Source: https://docs.vieneu.io/docs/integrations/mcp/troubleshooting ## Connecting | Symptom | Cause | Fix | |---|---|---| | "Couldn't reach the MCP server" / connection failed | Wrong URL | Use exactly `https://api.vieneu.io/mcp` — no trailing slash, no `/api` | | The VieNeu page says the request expired or was already handled | The sign-in page was left open more than 30 minutes, or approved twice | Go back to the app and connect again | | "Tài khoản chưa có gói token còn hạn" on the VieNeu page | No active plan | Buy or renew a plan on vieneu.io, then connect again | | "Tài khoản đã có 20 ứng dụng kết nối" | Too many connections | Remove old ones under **Developer → Gỡ kết nối** | | The VieNeu page signs in a different account than you expected | You were already signed in to vieneu.io in that browser | Click **Đổi tài khoản** on the page | | A red warning: the app "calls itself Claude/ChatGPT" but returns elsewhere | The request did not come from that app | Press **Từ chối** (Deny). Real Claude returns to `claude.ai`, ChatGPT to `chatgpt.com` | | A yellow note: the app runs on your computer | Normal for Claude Code, Cursor, VS Code and other desktop tools | Allow it if you just clicked connect in that tool | ## Using it | Symptom | Cause | Fix | |---|---|---| | The assistant says it has no VieNeu tools | The connector is off for this chat | Turn it on (Claude: **+ → Connectors**; ChatGPT: **+ → Developer mode**) | | New VieNeu features (audiobooks, "what can VieNeu do?") don't show up | The app keeps the list of tools it fetched when you connected | Disconnect VieNeu and connect again, then start a new chat (Claude: **Settings → Connectors → VieNeu → Disconnect**, then **Connect**). Apps that load their tools when they start (Claude Code, Cursor, VS Code) only need a restart | | Suddenly asked to sign in again | Removed on vieneu.io, or the sign-in expired | Reconnect from the app | | "Voice … is not available" | The assistant guessed a voice id | Ask it to search with `list_voices` first | | "Tài khoản không đủ quyền hoặc hết token…" | Out of tokens or plan expired | Top up or renew — retrying does not help | | "Audio đang được tạo…" with a job id | Long text still rendering | Ask the assistant to check again in a moment (it uses `get_speech_status`, free) | | No player, only a link | The app does not show MCP Apps (Claude Code and other CLIs), or you declined it | Use the link; in Claude choose **Allow** when asked to show the app | | The link stopped working | Links expire after 24 hours | Download it sooner, or find the audio in your **Library** on vieneu.io | ## Charges Every `text_to_speech` call is one synthesis, billed per character (minimum 50). Assistants sometimes offer to read the same sentence with a few different voices — each of those is charged. A failed synthesis is refunded automatically. See usage under **Developer → Usage** on vieneu.io. ## Still stuck? Contact VieNeu support from vieneu.io with the time of the attempt and the app you used. For developers: the server's discovery documents are at [`/.well-known/oauth-protected-resource`](https://api.vieneu.io/.well-known/oauth-protected-resource) and [`/.well-known/oauth-authorization-server`](https://api.vieneu.io/.well-known/oauth-authorization-server). --- # Cloud API overview Source: https://docs.vieneu.io/docs/cloud-api/overview The VieNeu Cloud API turns Vietnamese text into speech over HTTPS. You send text and a voice id, you get audio back — no model to download, no GPU to rent. :::note Two different products This section documents the **hosted API** at `api.vieneu.io`. The **SDK** section documents the separate on-device package, which runs a model on your own machine. They share voices and a name, and nothing else: different install, different billing, different code. If you are integrating VieNeu into a product, you almost certainly want this section. ::: If you would rather have audio in front of you before reading any of this, the **[Quickstart](./quickstart.md)** is a key, a voice and one call. ## The full reference Every endpoint, every field, every error is in the **[API reference](/api-reference)**, rendered from the OpenAPI specification the server itself generates. No login required. The same document is served as a plain file at **[`/openapi.json`](pathname:///openapi.json)** — point `openapi-generator`, `oapi-codegen`, Kiota or your language's equivalent at it and you have a typed client: ```bash curl -O https://docs.vieneu.io/openapi.json ``` It is generated from the API's own route definitions, and CI fails the build if the committed file no longer matches them — so the spec cannot quietly fall behind the *declarations* in the code. That is a narrower guarantee than it sounds, and worth stating plainly: the check compares the committed JSON against the decorators, not against what the server does. A decorator that describes an endpoint incorrectly ships a spec that is wrong and a check that is green — which is exactly how the cloning endpoint went on advertising an engine choice for months after the server started refusing one. Where this documentation and the spec disagree about behaviour, the pages here are the ones written against the running code. It covers `/api/v1` only — the application's own dashboard, billing and administration endpoints are not part of the public contract. ## Base URL ``` https://api.vieneu.io/api/v1 ``` ## Authentication Every request carries your API key, either way round: ```bash -H "Authorization: Bearer vn_sk_..." # or -H "X-API-Key: vn_sk_..." ``` Keys beginning `vn_sk_` are live; `vn_test_` keys work the same way but cap each request at 100 words. Create either on the **[Developer page](https://www.vieneu.io/#/developer?create=test)** in the VieNeu web app — the plaintext is shown once, at creation, and never again. A key is always issued against an **active token grant**, so a new account has to start one before the page will mint anything; without a grant, creation answers `403` with `No active API token grant found.` The Developer page offers a free 7-day trial that starts one, once per account. ## The two ways to synthesize **Synchronous** — you get the audio bytes in the response: ```bash curl https://api.vieneu.io/api/v1/audio/speech \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "input": "Xin chào, đây là VieNeu.", "voice": "Ngọc Lan" }' \ --output speech.mp3 ``` This is the [OpenAI-compatible endpoint](./openai-compatible.md) — the fastest way in if you already have an OpenAI client, and the one to use for short text. **Asynchronous** — for long text, submit a job and poll it: ```bash curl -X POST https://api.vieneu.io/api/v1/tts \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "text": "…", "voiceId": "Ngọc Lan" }' # → { "jobId": "…", "status": "queued" } # or, when that exact audio already exists (a cache hit): # { "jobId": "…", "status": "completed", "audioUrl": "https://…" } — nothing to poll curl https://api.vieneu.io/api/v1/tts/ \ -H "Authorization: Bearer $VIENEU_API_KEY" # → { "status": "completed", "audioUrl": "https://…" } ``` With several jobs in flight, poll them together — one request, up to 50 ids, each entry in the same shape as the single poll: ```bash curl "https://api.vieneu.io/api/v1/tts?ids=,," \ -H "Authorization: Bearer $VIENEU_API_KEY" # → { "jobs": [ { "jobId": "…", "status": "completed", "audioUrl": "…" }, … ], # "missing": [] } ``` Ids that are unknown or not yours land in `missing`; the batch is never refused because of one bad id. Prefer the batch form whenever you have more than one job queued: the edge proxy caps concurrent connections per source address, and a batch costs the server one database read where each single poll costs two. And when you need audio to start playing before the whole text is generated, use [streaming](./streaming.md). ### Retrying safely `POST /v1/tts` creates the job and charges the tokens **before** it answers. If your request times out, or the connection is reset while the answer is on its way — a `502` from the edge during a deploy looks exactly like this — you cannot tell "never happened" from "happened, answer lost", and a blind retry can pay twice. Send an `Idempotency-Key` header and it cannot. Choose the key yourself (8–128 printable ASCII characters, no whitespace; a UUID per job is the intended shape) and reuse it on every retry of *that* request: ```bash curl -X POST https://api.vieneu.io/api/v1/tts \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Idempotency-Key: 6f1c2e0a-7b3d-4c58-9e21-0a5d3b7f9c44" \ -H "Content-Type: application/json" \ -d '{ "text": "…", "voiceId": "Ngọc Lan" }' ``` - Same key, same body, first call finished → the **first** job comes back (same `jobId`, nothing charged again) with `Idempotent-Replayed: true`. - Same key, same body, first call still running → `409 IDEMPOTENCY_IN_FLIGHT` with `Retry-After: 1`. Retry with the **same** key. - Same key, different body → `422 IDEMPOTENCY_KEY_REUSED`. Keys are per request. - A first call that failed with a 4xx (quota, unknown voice) releases its key, so a corrected retry under the same key goes through. Keys are scoped to your API key and remembered for 24 hours. Every keyed response echoes the key back in `Idempotency-Key`. Without the header the route behaves exactly as it always has. ## Every operation Thirty-two operations on twenty-seven paths — the whole public surface. What each one does, and nothing about its fields: the request bodies live in the [API reference](/api-reference) alone, because a second description of the API is a second thing to keep true, and the one that goes stale is never the generated one. The reference deep-links by operation id, so `operationId` below is also how you jump straight to an entry — `/api-reference#operation/createTtsJob` — and how you name it to a generated client. **Synthesis** | Operation | `operationId` | What it does | |---|---|---| | `POST /v1/audio/speech` | `createSpeech` | OpenAI-compatible synthesis, sync or streaming | | `POST /v1/tts` | `createTtsJob` | Submit an asynchronous job | | `GET /v1/tts/{jobId}` | `getTtsJob` | Poll a job for status and a download URL | | `GET /v1/tts?ids=…` | `getTtsJobs` | Poll up to 50 jobs in one call | | `POST /v1/tts/stream` | `streamSpeech` | Low-latency framed streaming | | `POST /v1/dialogue` | `createDialogue` | Multi-speaker dialogue in one call | | `POST /v1/dub` | `createDub` | Re-voice an existing recording | | `POST /v1/srt` | `createSrtDub` | Dub a subtitle file into a timecode-aligned track | | `POST /v1/vapi/speech` | `vapiSpeech` | [Vapi](./vapi.md) custom-voice webhook — raw PCM for voice agents | **Voices and metadata** | Operation | `operationId` | What it does | |---|---|---| | `GET /v1/voices` | `listVoices` | The catalogue, plus your own cloned voices when authenticated. No key required | | `GET /v1/audio/voices` | `listOpenAiVoices` | Ids only, for OpenAI-compatible clients. **Does** need a key | | `GET /v1/engines` | `listEngines` | Live engines, features and billing multipliers. No key required | | `GET /v1/emotion-tags` | `listEmotionTags` | Reading styles and inline cue tags, per engine. No key required | **Usage** | Operation | `operationId` | What it does | |---|---|---| | `GET /v1/balance` | `getBalance` | Tokens this key can spend right now, the plan it bills and its daily / weekly caps. Not billed | | `GET /v1/usage` | `getUsage` | What this key (or your whole account) did over a window: calls, tokens, seconds of audio, per day / route / voice, errors | **Cloning — web only** Cloned voices are created in the web Studio at [vieneu.io/#/clone](https://vieneu.io/#/clone), not through this API. Enrolment is a guided job: the Studio denoises the clip, transcribes it, and lets you hear the result before the voice is saved. A bad reference degrades every later generation with that voice, not just one request, which is why the API no longer offers the shortcut. `POST /v1/voices`, `DELETE /v1/voices/{voiceId}`, `POST /v1/clone`, `POST /v1/upload` and `POST /v1/prepare` answer `410` with `code: "CLONE_WEB_ONLY"` and bill nothing. Using a clone through the API is unchanged: it appears in `GET /v1/voices` when you send your key (`"kind": "cloned"`), and its `clone_…` id is an ordinary `voiceId` on `POST /v1/tts` and `POST /v1/tts/stream`, at the same price as a catalogue voice. See [Cloned voices](./streaming.md#cloned-voices). **Webhooks** — see [Webhooks](./webhooks.md) for the event contract and a signature verifier you can paste in. | Operation | `operationId` | What it does | |---|---|---| | `POST /v1/webhooks` | `createWebhookEndpoint` | Register an endpoint. Returns the signing secret once | | `GET /v1/webhooks` | `listWebhookEndpoints` | List your endpoints | | `GET /v1/webhooks/{endpointId}` | `getWebhookEndpoint` | Read one endpoint | | `DELETE /v1/webhooks/{endpointId}` | `deleteWebhookEndpoint` | Delete an endpoint | | `GET /v1/webhooks/{endpointId}/deliveries` | `listWebhookDeliveries` | Recent delivery attempts, with the failure reason | | `POST /v1/webhooks/{endpointId}/rotate-secret` | `rotateWebhookSecret` | Rotate the signing secret, old one live for 24 hours | **Audiobooks** — a narrator and up to 30 character voices reading chapters of script. Books are made in the background, one chapter at a time, while the GPU fleet is idle: expect hours, not seconds. Books waiting together take turns, a chapter each. Each chapter is billed per character when it starts, refunded if it fails, mastered to MP3 and saved to the Library. Chapters are not reported as `tts.job.*` webhook events — follow the book. The same feature from a chat: [Audiobooks](../integrations/mcp/audiobooks.md). | Operation | `operationId` | What it does | |---|---|---| | `POST /v1/audiobooks` | `createAudiobook` | Create a draft: title, narrator, cast. LIVE keys only | | `GET /v1/audiobooks` | `listAudiobooks` | Your books, newest first | | `GET /v1/audiobooks/{bookId}` | `getAudiobook` | Progress, and a download link (24 h) per finished chapter | | `POST /v1/audiobooks/{bookId}/chapters` | `addAudiobookChapter` | Add a chapter as a script, or continue one not started (`appendTo`) | | `POST /v1/audiobooks/{bookId}/start` | `startAudiobook` | Queue it; `402` when the plan cannot pay. Also resumes and redoes failed chapters | | `POST /v1/audiobooks/{bookId}/cancel` | `cancelAudiobook` | Stop. Chapters not started are never billed | ### Voice ids differ by engine A voice belongs to one engine, and since **2026-09-24** the cloud API has exactly one: `v4`, whose ids are Vietnamese display names — `Ngọc Lan`, not a slug. The retired `v3` catalogue used opaque `vieneu-…` slugs, and the two catalogues shared almost no ids, so a voice copied from an old `v3` listing does not resolve anywhere any more. It is rejected with `400` rather than silently mapped to a `v4` voice — and that is now the most common cause of `400`s on the synthesis routes for callers who hard-coded a voice before the retirement, shipped, and never had a reason to look at the id again. **Resolve ids at runtime, scoped to the engine you are about to use:** ```bash curl -s "https://api.vieneu.io/api/v1/voices?engine=v4" ``` No key needed. `GET /v1/voices` with no filter lists the same `v4` catalogue (plus your own `v4` clones when you send a key); `?engine=v3` answers `400` with the retirement message quoted under [Engines](#engines). ## Engines Requests carry an optional `engine`. Since **2026-09-24 08:00 (GMT+7)** the only value the cloud API accepts is `v4`, which is also the default, so the field can simply be omitted. `v3` was retired that morning: any request naming it — on `POST /v1/tts`, `/v1/tts/stream`, `/v1/audio/speech`, `/v1/dialogue`, `/v1/dub`, `/v1/srt`, or as `GET /v1/voices?engine=v3` — is refused with `400` before anything is billed, and the message says why: ``` Engine "v3" was retired on 2026-09-24. Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues share no ids. ``` If you had pinned `v3`, the fix has three parts: send `"engine": "v4"` (or nothing), take a voice id from `GET /v1/voices?engine=v4` rather than reusing the old one, and re-enrol any clone you made on `v3` in the [Studio](https://vieneu.io/#/clone). Audio you already generated on `v3` is untouched — library entries and download links keep working. The engine is billed at its own multiplier, applied on top of the per-character rate: | Engine | Multiplier | |---|---| | `v4` | 3× | That figure is set per deployment and can change, so the table above is a snapshot, not a contract. **`GET /v1/engines` returns the live values** — no API key needed. If you are budgeting a large workload, read them from there rather than from this page: ```bash curl -s https://api.vieneu.io/api/v1/engines ``` ```json { "globalMultiplier": 1.3, "engines": [ { "key": "v4", "isDefault": true, "sampleRate": 48000, "billingMultiplier": 3, "features": ["generate", "clone", "dialogue", "stream", "dub", "srt"] } ] } ``` ### The second multiplier `globalMultiplier` in that response is a platform-wide lever sitting **on top of** the per-engine rate. The full charge for one request is: ``` tokens = characters × engine multiplier × globalMultiplier ``` It spends long stretches at `1`, which is exactly why it is easy to miss — and why it is returned alongside the engine rate rather than documented somewhere else. At the time of writing it is `1.3`. Budget from the per-engine number alone and your estimate is short by precisely this factor on any deployment where it has been moved. (A third factor, the AI text-refinement surcharge, applies only to requests that ask for that step; it is not part of the base rate.) `features` is also how you check, in code, what an engine can do — `stream` is on `v4`'s list, so every stream runs there and is billed at the `v4` rate, the same rate as every other call now. A voice belongs to exactly one engine — `GET /v1/voices?engine=v4` lists that engine's catalogue. Passing a voice from the retired `v3` catalogue is rejected rather than silently substituted. **Streaming is `v4` only**, and was before the retirement, for a reason worth keeping in mind if you ever compare providers: `v3` never really streamed — it delivered its chunks in a burst at the end — so `POST /v1/tts/stream` refused it rather than making a time-to-first-audio promise it could not keep. Now that `v3` is gone from every route the gate is moot: a stream request naming `v3` gets the same retirement `400` as any other request, and so does `POST /v1/audio/speech` with `"model": "vieneu-v3"`, which until 2026-09-24 accepted the request, answered 200, and delivered the burst at the end with nothing in the response saying so. If you are pinning an engine for a streaming client, pin `v4` — see [Drop-in for OpenAI-compatible apps](../integrations/openai-clients.md). A cloned voice streams too, provided it is enrolled on `v4` — every clone created since 2026-08-28 is, and an older `v3`-enrolled clone can no longer be rendered at all until it is re-enrolled. See [Cloned voices](./streaming.md#cloned-voices). ## Billing Synthesis is billed per **submitted character**, so the cost of a call is predictable from the request itself, before you make it. Minimum 50 characters per request. A request that fails to produce audio is refunded automatically. `aiRefine` defaults to **off** on the API. With it off your text is synthesized exactly as sent. Turn it on (`"aiRefine": true`) for the AI pass the web app uses — formulas, acronyms and mixed-in English read correctly, and the content is checked — billed with a surcharge and one extra round-trip of latency. Deterministic text preparation runs either way, so Vietnamese is pronounced correctly regardless. ### The one endpoint that is not billed per character `POST /v1/dub` is priced by the **output audio duration** — 65 tokens per second, with a 150-token floor — and charged **only on success**. Dubbing runs speech recognition before it runs synthesis, so the work tracks how long the recording is, not how many characters the speaker happened to fit into it. The engine multiplier still applies on top. That makes dub the one call whose cost you cannot compute from the request body before sending it. An upload is gated for affordability against its own duration first, so a request you cannot pay for is refused before any GPU time is spent on it. ### Seeing what you were billed for `GET /v1/balance` is the quick question — *can I afford the next call?* `remaining` is what one call can spend right now: the key's plan's tokens, or less when a daily or weekly cap is lower. `plan` has the totals, each cap's remainder and reset time (`null` when there is no such cap) and the expiry; `totalRemainingTokens` adds up every active plan on the account. A synthesis needs at least `50 × engine multiplier` tokens. `GET /v1/usage` answers "am I being counted correctly" without a support ticket. It reads the same per-call ledger the operators see: for the calling key (default) or every key on the account (`scope=account`), over a window of up to 92 days (`from`, `to`; default the last 30 days), it returns totals — calls by outcome, tokens, characters, seconds of audio, average and p95 generation time — and breakdowns per day, per route, per voice and per error code. A few things to know when reading it. `POST /v1/tts` rows start as `queued` and become `ok` or `error` when the job finishes; a request served from the content cache is `cache_hit` (billed, no generation time). `tokenCost` is what was deducted **before** the call ran; `refunded: true` on a row means it was given back. A call refused by the rate limiter never reaches the ledger — count those from your own 429s. `requestCount` on your key's own record is something else: how many times the key *authenticated*, polling included. ## Rate limits and request ids There are **two** limiters in front of you, and they count different things. An integration that only knows about one of them will misread half its 429s. **The application limiter counts your API key.** This is the one that matches what you bought: it keys on the key itself, so spreading calls across machines does not buy you more and sharing an office network does not cost you any. Limits are per route, and every route on `/v1` sets its own: | Routes | Per minute | |---|---| | `/v1/audio/speech`, `/v1/tts`, `/v1/tts/stream`, `/v1/vapi/speech` | 300 | | `/v1/dialogue` | 20 | | `/v1/dub`, `/v1/srt` | 10 | | Webhook management | 10–30, per operation | | `GET /v1/voices`, `/v1/audio/voices`, `/v1/engines`, `/v1/emotion-tags`, `GET /v1/tts/{jobId}`, `GET /v1/tts?ids=` | not throttled | When this limiter refuses you, it says so: ``` X-RateLimit-Limit: 300 X-RateLimit-Remaining: 0 X-RateLimit-Reset: 43 Retry-After: 43 X-Request-Id: 0f7c… ``` `X-RateLimit-Reset` is when the window rolls; `Retry-After` is when you may call again. They usually agree, and when they do not, honour `Retry-After`. **The edge proxy counts your source address.** In front of the application sits nginx, which can only key on an IP. On `/api/v1` it allows **3000 requests per minute and 300 concurrent connections per address**. That is deliberately loose — it is DDoS padding, not a product limit — but it is real, it is **shared with every other caller behind your address**, and it knows nothing about your key. One customer calling from a dozen machines looks like a dozen customers there, and a dozen customers behind one integration platform's egress IP look like one. **The tell is the headers.** A `429` carrying `X-RateLimit-*` came from the application and is about your key. A bare `429` or `503` with none of those headers and no `Retry-After` came from the edge and is about your IP — nginx answers with its own page and none of our headers. That difference decides what to do: back off per key in the first case; in the second, back off *and* look at how many machines share that egress address, because widening concurrency will make it worse. Streaming is the case where this matters most, because `/v1/tts/stream` matches a different nginx location than the rest of `/api/v1` and inherits a much tighter connection cap — see [Limits](./streaming.md#limits). Every response, success or failure, carries `X-Request-Id`. Quote it in a support request — it is the id our logs are keyed by. ## Errors Native endpoints return the usual shape: ```json { "statusCode": 403, "message": "Insufficient tokens", "traceId": "0f7c…" } ``` `/v1/audio/speech` returns OpenAI's shape instead, so OpenAI clients can parse it. | Status | Meaning | |---|---| | 400 | Malformed request — an unknown voice, a bad format, text the validator refused | | 401 | Missing, malformed or revoked API key | | 403 | Out of tokens, grant expired, or your plan does not include this engine | | 422 | Content refused by moderation (only when `aiRefine` is on) | | 429 | Rate limit or token quota — see above for telling them apart | | 503 | No worker available for the requested engine or format; retry shortly | Which refusals carry a machine-readable `code`, which do not, and what to do about each is on the [Errors](./errors.md) page. --- # Quickstart Source: https://docs.vieneu.io/docs/cloud-api/quickstart Three steps, and the third one returns audio. Nothing here needs an SDK. ## 1. Get an API key Keys live on the **[Developer page](https://www.vieneu.io/#/developer?create=test)** in the VieNeu web app. Create one, copy it once — the plaintext is shown at creation and never again — and put it in your environment: ```bash export VIENEU_API_KEY="vn_sk_..." ``` A key can only be issued against an **active token grant**, so a brand-new account has to start one first; the Developer page offers a free 7-day trial that does exactly that, once per account. Without a grant, key creation answers `403` with `No active API token grant found.` Keys beginning `vn_sk_` are live. `vn_test_` keys behave identically — same billing, same limits, same endpoints — except that each request is capped at 100 words, which makes them safe to paste into a sample repository. ## 2. Pick a voice The catalogue is public — this call needs no key: ```bash curl -s "https://api.vieneu.io/api/v1/voices?engine=v4" | head -c 400 ``` ```json {"voices":[{"id":"Ngọc Lan","description":"Giọng nữ, giọng trầm dịu dàng", "name":"Ngọc Lan","gender":"female","region":"south","engine":"v4", "kind":"catalog"}, …]} ``` The `id` field is what you send. **Always filter by engine.** A voice belongs to one engine, the two catalogues are only partly interchangeable, and sending an id from the wrong one is a `400` rather than a substitution — see [voice ids differ by engine](./overview.md#voice-ids-differ-by-engine). ## 3. Synthesize `POST /v1/audio/speech` returns the audio bytes in the response — no job, no polling. It is OpenAI's endpoint shape, so an OpenAI client works against it unchanged; see [OpenAI-compatible endpoint](./openai-compatible.md). ### curl ```bash curl https://api.vieneu.io/api/v1/audio/speech \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "input": "Xin chào, đây là VieNeu.", "voice": "Ngọc Lan" }' \ --output speech.mp3 ``` ### Python ```python # pip install requests import os import requests resp = requests.post( "https://api.vieneu.io/api/v1/audio/speech", headers={"Authorization": f"Bearer {os.environ['VIENEU_API_KEY']}"}, json={"input": "Xin chào, đây là VieNeu.", "voice": "Ngọc Lan"}, timeout=120, ) resp.raise_for_status() with open("speech.mp3", "wb") as f: f.write(resp.content) print("wrote speech.mp3", len(resp.content), "bytes") ``` ### JavaScript ```js // Node 18+ — no dependencies. import { writeFile } from 'node:fs/promises'; const resp = await fetch('https://api.vieneu.io/api/v1/audio/speech', { method: 'POST', headers: { Authorization: `Bearer ${process.env.VIENEU_API_KEY}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ input: 'Xin chào, đây là VieNeu.', voice: 'Ngọc Lan' }), }); if (!resp.ok) throw new Error(`HTTP ${resp.status}: ${await resp.text()}`); await writeFile('speech.mp3', Buffer.from(await resp.arrayBuffer())); console.log('wrote speech.mp3'); ``` You now have an mp3. `response_format` changes that — `wav`, `opus`, `pcm` and 8 kHz `ulaw` are all available, with the field-level detail in the [API reference](/api-reference) under `createSpeech`. ## Where to go next - **Long text** — `/v1/audio/speech` holds the connection open for the whole synthesis. Past a few paragraphs, submit a job instead: see [the two ways to synthesize](./overview.md#the-two-ways-to-synthesize). - **Playback before the text finishes** — [Streaming](./streaming.md). - **When a call fails** — [Errors](./errors.md) lists every machine-readable `code` and what to do about it. - **Everything else** — the [API reference](/api-reference) is generated from the server's own route definitions and covers every field of every endpoint. --- # OpenAI-compatible TTS endpoint Source: https://docs.vieneu.io/docs/cloud-api/openai-compatible VieNeu's native public API is asynchronous (`POST /v1/tts` returns a `jobId` you poll). `POST /v1/audio/speech` is a **drop-in for OpenAI's endpoint of the same name**: point an OpenAI SDK at VieNeu's base URL, give it your VieNeu key, and it works without any other change. This matters more than it looks. A large amount of software already speaks this one endpoint — Open WebUI, SillyTavern, LobeChat, LiteLLM, LiveKit's and Pipecat's OpenAI plugins, a long tail of scripts — and all of them accept a custom `base_url`. Supporting this shape is what makes VieNeu usable in them with no plugin, no adapter and no work on our side. ## Endpoint ``` POST /api/v1/audio/speech Authorization: Bearer # vn_sk_... (X-API-Key also accepted) Content-Type: application/json ``` Returns the audio **bytes synchronously** — no polling. ## Request body | Field | OpenAI | VieNeu behavior | |-------|--------|-----------------| | `input` | required | the text to synthesize ✅ | | `model` | required | selects the engine — and since 2026-09-24 there is one, `v4`. OpenAI's names (`tts-1`, `tts-1-hd`, `gpt-4o-mini-tts`), `vieneu` and `vieneu-v4` all resolve to it. Any other value is accepted and ignored, as before — except a `vieneu-…` name for an engine that does not exist or has been retired, which is rejected: `vieneu-v3` answers 400 with `Model 'vieneu-v3' is retired. Engine "v3" was retired on 2026-09-24. …`. | | `voice` | `alloy`, … | a **VieNeu** voice id — list them with `GET /v1/audio/voices`. OpenAI's voice names are not mapped. Omit for the default voice. | | `response_format` | `mp3` (default) | `mp3` (default), `wav`, `opus`, `pcm`, plus VieNeu's `ulaw`. `aac` and `flac` return 400. | | `speed` | 0.25–4.0 | accepted across OpenAI's full range and **clamped** to 0.5–2.0, where the engine holds quality. A valid OpenAI value never returns an error. | | `stream_format` | `audio` \| `sse` | supported for `pcm` and `ulaw` — see [Streaming](#streaming). | | `instructions` | (gpt-4o-mini-tts) | accepted and ignored. Use inline cue tags instead. | | `sample_rate` | — | VieNeu extension: 8000, 16000, 22050, 24000, 44100 or 48000. | | `emotion` | — | VieNeu extension left over from `v3`: `natural` (default) or `storytelling`. Accepted and **ignored** on `v4` — see below. | | `aiRefine` | — | VieNeu extension, **default `false`** — see [Billing](#billing). | :::caution `emotion` does nothing any more `emotion` was a `v3` parameter, and `v3` was retired from the cloud API on 2026-09-24. `v4` — now the only engine — has **no reading styles**. It renders three inline cue tags — `[cười]`, `[thở dài]` and `[hắng giọng]` — and **deletes every other tag from your text** rather than voicing it. A request that sets `"emotion": "storytelling"` is accepted, billed, and read in the ordinary voice; it is not an error, so nothing in the response tells you the field went nowhere. Naming `"model": "vieneu-v3"` to get the old behaviour back is a 400. `GET /v1/emotion-tags` returns what the engine can actually render, and it needs no API key. With no `engine` it answers for the default — `v4` — so the reply is the three cues above and an empty list of styles; `?engine=v3` is a 400. Ask it rather than assuming, especially if you are carrying a tag list written against `v3`. ::: ### Formats `pcm` and `ulaw` are **headerless**: the bytes carry no sample rate, so read it from the `X-Sample-Rate` response header. Rates: `pcm` defaults to **24000**, matching what OpenAI documents so a client following their contract plays it at the right speed. Everything else defaults to 48000. `opus` is always 48 kHz (Opus itself is) and `ulaw` always 8 kHz, which is what a phone line wants — passing a conflicting `sample_rate` is rejected rather than quietly ignored. > mp3 and opus need `ffmpeg` on the worker that serves the request. A node that > predates it answers **503** naming the format, rather than returning something > that is not the format you asked for. The one exception is a request that never > named a format: since mp3 is *our* default rather than your choice, those fall > back to wav instead of failing. ## Examples **curl** ```bash curl https://api.vieneu.io/api/v1/audio/speech \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu.", "voice": "Ngọc Lan" }' \ --output speech.mp3 ``` **OpenAI Python SDK** (unmodified, pointed at VieNeu) ```python from openai import OpenAI client = OpenAI(api_key="vn_sk_...", base_url="https://api.vieneu.io/api/v1") resp = client.audio.speech.create( model="tts-1", voice="Ngọc Lan", # a VieNeu voice id, from GET /v1/audio/voices input="Xin chào, đây là VieNeu.", ) resp.stream_to_file("speech.mp3") ``` **Telephony (8 kHz mu-law)** ```bash curl https://api.vieneu.io/api/v1/audio/speech \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "input": "Xin chào", "response_format": "ulaw" }' \ --output line.raw # raw G.711 mu-law, 8000 Hz, no header ``` ## Streaming Set `stream_format` to start receiving audio while it is still being generated, instead of waiting for the whole file. - **`audio`** — the bytes arrive as chunked transfer encoding. Append them; the result is the same audio you would have got in one piece. - **`sse`** — Server-Sent Events. `speech.audio.delta` events carry base64 audio; the stream ends with exactly one `speech.audio.done` (the audio is complete) or `speech.audio.error` (it is not). ```python with client.audio.speech.with_streaming_response.create( model="tts-1", voice="Ngọc Lan", input="…", response_format="pcm", extra_body={"stream_format": "audio"}, ) as resp: resp.stream_to_file("speech.pcm") # raw s16le, 24 kHz ``` **`speech.audio.done` is the only proof the audio is whole.** If a worker dies mid-generation the stream stops, and a truncated stream is otherwise indistinguishable from a short one. When that event does not arrive, discard the audio — the request is refunded automatically. ### Streaming is `pcm` and `ulaw` only Streamed audio is generated in chunks and each chunk is encoded independently, so joining them only produces a valid result for the headerless formats: | Format | Streamable | Why not | |---|---|---| | `pcm`, `ulaw` | ✅ | raw samples; concatenation *is* playback | | `wav` | ❌ | every chunk repeats its 44-byte header mid-file | | `mp3` | ❌ | each chunk re-applies the encoder's delay — measured at ~24 ms of inserted silence per seam, plus a click | | `opus` | ❌ | each chunk is a complete Ogg stream; most browsers play only the first one | Asking for a non-streamable format with `stream_format` returns 400 rather than shipping audio with gaps in it. For a complete mp3 or opus **file**, drop `stream_format` — the synchronous response has none of these problems. This will widen once the worker can hold a single encoder open across a whole stream; today it starts a fresh one per chunk. ## Voices ``` GET /api/v1/audio/voices[?engine=v4] → { "voices": ["Ngọc Lan", …] } ``` A companion to this endpoint for OpenAI-compatible clients, which look for this route to fill their voice picker. `GET /v1/voices` is the richer version — names, gender, region, and your own cloned voices. ## Errors Returned in OpenAI's shape: ```json { "error": { "message": "...", "type": "invalid_request_error", "param": "input", "code": null } } ``` - `400 invalid_request_error` — missing `input`, an unknown or retired `vieneu-…` model (`vieneu-v3`, since 2026-09-24), an unsupported `response_format`, `stream_format` with a format that cannot be streamed, or a `sample_rate` that contradicts the format. - `401 authentication_error` — the key is missing, malformed or revoked. - `403 insufficient_quota` — the key's token grant is exhausted or expired. **Not 402:** nothing on `/v1` answers `402` at all. Being out of credit splits across two statuses here, and the split is the part worth branching on: an exhausted or expired grant is `403 insufficient_quota` and **will not clear on its own**, while a spent daily or weekly cap is `429 rate_limit_exceeded` and will. So `insufficient_quota` means top up or renew; `rate_limit_exceeded` means wait. - `422 content_policy_violation` — refused by moderation (only when `aiRefine` is on). - `429 rate_limit_exceeded` — the rate limiter throttled you, **or** a daily or weekly token cap is spent. This envelope cannot tell those two apart: the native routes name the cap in `code` and say when it lifts in `resetAt`, and neither field survives the translation into OpenAI's shape. See [Errors](./errors.md#the-openai-compatible-route). - `503` — no worker for the requested engine, or none that can encode the requested format. - `500 api_error` — synthesis failed (the token charge is **refunded**). One caveat on `403`. The quota `403` above is typed by the handler itself. A `403` raised *before* the handler runs — a guard refusing a feature your plan does not include — is typed from the status alone, which folds `401` and `403` together as `authentication_error`. So `insufficient_quota` always means money, but `authentication_error` can mean either the key or the plan. Every response carries `X-Request-Id`; quote it in a support request. The body does not repeat it — OpenAI's envelope has no field for it, so read the header. ## Billing Tokens are deducted from the API key's grant by the **submitted** character count, so the cost of a call is predictable from the request alone. A failed synthesis is refunded automatically. `aiRefine` defaults to **`false`** here, unlike the web app, where it is on. With it off the text is synthesized as submitted: no AI moderation, no pronunciation normalization, no surcharge, and one less model round-trip of latency. Set it to `true` to get the web app's behaviour — formulas, acronyms and mixed-in English read correctly, content checked — billed with the AI surcharge. Deterministic text preparation (chemistry spelling, ALL-CAPS folding, the sea-g2p phoneme layer) runs either way. `aiRefine` controls only the AI call. ## Differences from OpenAI - `voice` expects a VieNeu voice id, not an OpenAI voice name. - `aac` and `flac` are not supported; `ulaw` and `sample_rate` are additions. - `stream_format` covers `pcm` and `ulaw` only, not every format. - `speed` outside 0.5–2.0 is clamped rather than honoured exactly. - The richer VieNeu features — multi-speaker `dialogue`, `dub`, SRT dubbing — have no OpenAI equivalent. Use the native `/v1/*` endpoints. Cloned voices are made in the web Studio and then used as an ordinary `voiceId` on `POST /v1/tts` (not on this route — a `clone_…` id here is a 400). --- # Streaming Source: https://docs.vieneu.io/docs/cloud-api/streaming Long text takes as long to synthesize as it takes. Streaming lets playback start on the first sentence — typically **1–2 seconds** on the standard plans, or a ceiling agreed per contract on Enterprise — instead of after the last one. :::info Engine — and what it costs Streaming runs on **`v4`**. It always did — `v3` delivered its chunks in a burst at the end, so a stream request naming it was refused rather than served under a time-to-first-audio promise it could not keep — and since **2026-09-24** `v4` is the only engine on the cloud API at all, so there is nothing to choose: omit `engine`, or send `v4`. Naming `v3` on this route now gets the same `400` as on every other: `Engine "v3" was retired on 2026-09-24. Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues share no ids.` **The price is the `v4` price.** A stream is billed per submitted character at `v4`'s multiplier — today **3×**, times the platform-wide `globalMultiplier` — exactly like a `POST /v1/tts` job for the same text. The old comparison, a stream at `v4`'s rate against a job on `v3`'s 1.5×, no longer means anything, because the 1.5× rate is gone from every route; if you sized a budget against it, resize it. `GET /v1/engines` returns the live multipliers, no key required — see [Engines](./overview.md#engines). This applies to cloned voices as well. Clones have enrolled on `v4` since 2026-08-28, so they stream like any other `v4` voice; an older `v3`-enrolled clone has to be re-enrolled first. See [Cloned voices](#cloned-voices). ::: There are two streaming endpoints, and which you want depends on who is reading the bytes: - **`POST /v1/audio/speech` with `stream_format`** — plain audio or SSE, and what an OpenAI client already knows how to read. Start here. - **`POST /v1/tts/stream`** — VieNeu's native framed stream. Use it when you want each chunk delivered as a separately decodable unit, or when you want a positive signal that the stream finished. ## The native stream ```bash curl -N -X POST https://api.vieneu.io/api/v1/tts/stream \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "text": "…", "voiceId": "Ngọc Lan" }' \ --output stream.bin ``` The response body is a sequence of length-prefixed frames: ``` [4-byte big-endian uint32 = N][N bytes of audio] … repeated … [4-byte big-endian uint32 = 0] ← end of stream ``` Read a length, read that many bytes, repeat. By default each payload is a **self-contained WAV**, so you can hand a frame straight to a decoder without waiting for the rest. Response headers tell you what you actually got: | Header | Meaning | |---|---| | `X-Sample-Rate` | Sample rate of the audio, in Hz | | `X-Output-Format` | The encoding: `wav`, `mp3`, `opus`, `pcm` or `ulaw` | | `X-Stream-Format` | `len32-wav-chunks`, or `len32-frames` for other encodings | ### The zero-length frame is the point A stream that finishes sends a final frame with length `0`. A stream cut short by a failure mid-generation simply **stops**, with no such frame. Its absence means the stream was cut short: discard the audio. You are not billed for a stream that never sent it — the refund is automatic. **But the marker alone is not success.** A stream can arrive complete and still carry no speech — a generation that produced nothing sends its heartbeats (see below) and then a clean end-of-stream. We do not bill those either, so a caller who treats the marker as sufficient will book a success, save a silent file, and end the month reconciling against an invoice that never charged for it. The condition we bill on, and the one you should use, is both halves: > **the end-of-stream marker arrived, AND at least one frame carried samples.** ### Heartbeat frames A stream that goes quiet for a while sends a **heartbeat**: a valid WAV file containing **zero samples**, 44 bytes on the wire. Its only job is to keep the connection from being closed for inactivity by whatever sits between us — a proxy, a corporate gateway, a mobile carrier's NAT. Heartbeats appear **only on WAV streams** (`X-Stream-Format: len32-wav-chunks`), at most one every `X-Stream-Heartbeat` seconds, and only while the synthesizer has produced nothing new. A stream that flows normally never sends one. **You must skip them.** The trap is that a heartbeat is a *well-formed* WAV, so a decoder will not reject it — `decodeAudioData` returns a zero-length buffer rather than throwing, and writing the frame to a file leaves a stray 44-byte header in the middle of your audio. Checking that a frame is non-empty is not enough; ask whether it carries samples: ```python def has_samples(frame: bytes) -> bool: """False for a heartbeat: a valid WAV whose `data` chunk is empty.""" if not frame: return False # nothing there at all if len(frame) < 12 or frame[:4] != b"RIFF" or frame[8:12] != b"WAVE": return True # not a WAV we can read — assume audio off = 12 while off + 8 <= len(frame): chunk_id = frame[off:off + 4] size = int.from_bytes(frame[off + 4:off + 8], "little") if chunk_id == b"data": return size > 0 off += 8 + size + (size % 2) # RIFF chunks are word-aligned return True ``` Walk the chunks rather than assuming the `data` chunk starts at byte 36 — a 44-byte header is the common case, not a rule. When you cannot read the header, treat the frame as audio. Guessing wrong in that direction costs you one odd frame; guessing wrong in the other throws away speech your listener was waiting for. Raw codec streams (`pcm`, `ulaw`, `mp3`, `opus` — `X-Stream-Format: len32-frames`) never carry heartbeats, because injecting a fake WAV into a codec bitstream would be injecting garbage. If you stream those, every frame is audio. ### Formats Pass `outputFormat` (and optionally `sampleRate`) to change the payload encoding: ```json { "text": "…", "voiceId": "Ngọc Lan", "outputFormat": "ulaw" } ``` | Format | Notes | |---|---| | `wav` | Default. Each frame is a self-contained file, 48 kHz. | | `pcm` | Raw signed 16-bit little-endian, **no header** — read the rate from `X-Sample-Rate`. | | `ulaw` | Raw G.711 mu-law, always 8 kHz. What a phone line wants. | | `mp3`, `opus` | Available, but see the warning below. | Valid `sampleRate` values are 8000, 16000, 22050, 24000, 44100 and 48000. The framing never changes, whatever the encoding — so the zero-length terminator means the same thing in all of them. :::warning mp3 and opus frames do not join cleanly Every frame is encoded as a standalone file, so **decode each frame separately** — do not concatenate them. Concatenated mp3 gains about 24 ms of silence at each seam (the encoder delay, re-applied per frame) plus an audible click. Concatenated opus is a chain of complete Ogg streams: ffmpeg reads it, most browsers stop at the first frame. If you want one continuous mp3 or opus body rather than frames, request the whole file without streaming — the synchronous path encodes it in one pass and has none of these artifacts. `/v1/audio/speech`'s `stream_format` therefore accepts only `pcm` and `ulaw`. ::: ## How a stream ends Five outcomes, and they are not all failures. Your integration should tell them apart, because three of them mean "try again" and two do not. | What you see | What it means | Billed? | |---|---|---| | Frames, then a zero-length frame, at least one frame carrying samples | Success | **yes** | | Frames, then a zero-length frame, but every frame was a heartbeat | Synthesis produced nothing | no — refunded | | The body just stops, no zero-length frame | Cut short mid-generation | no — refunded | | `503` with `"code": "STREAM_BUSY"` | Every node is healthy but streaming capacity is used up | no — nothing was charged | | `502` | Every node for this engine failed | no — refunded | `STREAM_BUSY` carries `"fallback": "generate"`, and that is a real instruction: the queued `POST /v1/tts` path has separate capacity and will accept the work right now. Retrying the stream immediately usually will not. A stream that reaches us but produces no audio for **120 seconds** is abandoned server-side and ends without the zero-length frame — so it arrives as the "cut short" row above. That ceiling exists so a wedged node cannot hold your connection open indefinitely. ### If your client disconnects **You are charged.** Closing the connection part-way — a user pressing Stop, a timeout on your side, a crashed worker of your own — bills the request. You received audio; we generated it. This is deliberate, and it is the one case where "no zero-length frame" does not mean a refund. If you retry after aborting, budget for both attempts. ### Limits | Limit | Value | On exceeding | |---|---|---| | Concurrent streams | 4 | `429`, `code: STREAM_CONCURRENCY` | | Synthesis requests per minute | 300 | `429`, honour `Retry-After` | | Text length | 50 000 characters | `400` | Concurrency is counted against the **token grant** behind your key, not the key itself, so several keys issued on one account normally share the four rather than each getting four of their own. The slot is taken before the balance is touched: a stream refused here costs nothing. :::caution The stream route has a much tighter connection cap `POST /v1/tts/stream` is matched by its own nginx location — the one that turns response buffering off, without which the whole point of streaming is lost. That location does not inherit `/api/v1`'s generous per-address connection allowance of 300; it falls back to the server-wide one of **20 concurrent connections per source IP**, and to a per-address request rate of 500/min rather than 3000. A stream holds its connection open for the whole synthesis, so those 20 are held, not cycled. Twenty concurrent streams from one egress address is a realistic number for a busy integration, and the twenty-first is refused by nginx with a bare `503` — an HTML error page, no `Retry-After`, no `X-RateLimit-*`, and none of the JSON a `STREAM_BUSY` carries. That absence is how you tell the two `503`s apart: ours has a body and a `"fallback"` you can act on, the edge's has neither. It is also counted across everyone behind that address, not just you. See [Rate limits](./overview.md#rate-limits-and-request-ids). ::: ## Cloned voices Pass a `clone_…` id as `voiceId` and it streams like any other voice — the same frames, the same terminator, the same headers: ```json { "text": "…", "voiceId": "clone_9f2c1e04-…" } ``` The ids come from `GET /v1/voices` with your API key (they are listed alongside the catalogue, tagged `"kind": "cloned"`). You create the voice itself in the web Studio at [vieneu.io/#/clone](https://vieneu.io/#/clone) — the API no longer enrols voices — and it is usable here the moment it is saved. Both your own clones and any an administrator has published to the catalogue work. **It costs the same as a preset.** Streaming is billed per submitted character × the engine's multiplier, and cloning adds no multiplier of its own. Enrolment is charged in the Studio, not here, so a stream with a `clone_…` voice costs exactly what the same text costs with a catalogue voice. :::caution Clones enrolled before 2026-08-28 Every clone created since 2026-08-28 enrols on `v4`, and those stream exactly as described above. A clone enrolled earlier was cut for `v3`, and `v3` was retired from the cloud API on 2026-09-24: such a voice is no longer listed by `GET /v1/voices` and cannot be rendered through the API on any route — the only engine that could read its reference clip is the one that now answers `400`. Re-enrol it in the [Studio](https://vieneu.io/#/clone); the new `clone_…` id streams like any other `v4` voice. Audio you generated with the old voice before the retirement is unaffected. ::: Two failures are worth distinguishing, and both arrive as `400` **before anything is billed**: | Message | What happened | |---|---| | `Cloned voice "…" was not found among your voices.` | Wrong id, or a clone belonging to another account | | `Cloned voice "…" is no longer available.` | The voice exists but its reference clip is gone — usually deleted mid-request | ## Reference decoder — Python ```python import struct, requests def stream_frames(text, voice, api_key, output_format="wav"): """Yield each audio frame. Raises if the stream was truncated.""" resp = requests.post( "https://api.vieneu.io/api/v1/tts/stream", headers={"Authorization": f"Bearer {api_key}"}, json={"text": text, "voiceId": voice, "outputFormat": output_format}, stream=True, ) resp.raise_for_status() print("sample rate:", resp.headers.get("X-Sample-Rate")) buf, complete, any_audio = bytearray(), False, False for chunk in resp.iter_content(chunk_size=8192): buf.extend(chunk) # A frame may span chunks, and several may arrive in one. while len(buf) >= 4: (length,) = struct.unpack(">I", buf[:4]) if length == 0: # end-of-stream marker complete = True del buf[:4] continue if len(buf) < 4 + length: # frame not all here yet break frame = bytes(buf[4:4 + length]) del buf[:4 + length] if has_samples(frame): # skip heartbeats — see above any_audio = True yield frame if not complete: raise RuntimeError("stream truncated — discard this audio") if not any_audio: # Kết thúc sạch nhưng không một mẫu nào: chúng tôi cũng không tính tiền # ca này. Coi nó là thành công là ghi sổ lệch với hoá đơn. raise RuntimeError("stream carried no audio — not billed, do not save") ``` ## Reference decoder — JavaScript ```js /** False for a heartbeat: a valid WAV whose `data` chunk is empty. */ function hasSamples(frame) { if (frame.length === 0) return false; // nothing there at all if (frame.length < 12) return true; // too short to read — assume audio const dv = new DataView(frame.buffer, frame.byteOffset, frame.byteLength); const tag = (o) => String.fromCharCode(...frame.subarray(o, o + 4)); if (tag(0) !== 'RIFF' || tag(8) !== 'WAVE') return true; // not WAV — assume audio let off = 12; while (off + 8 <= frame.length) { const size = dv.getUint32(off + 4, true); if (tag(off) === 'data') return size > 0; off += 8 + size + (size % 2); // RIFF chunks are word-aligned } return true; } async function* streamFrames(text, voice, apiKey, outputFormat = 'wav') { const resp = await fetch('https://api.vieneu.io/api/v1/tts/stream', { method: 'POST', headers: { Authorization: `Bearer ${apiKey}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ text, voiceId: voice, outputFormat }), }); if (!resp.ok) throw new Error(`HTTP ${resp.status}`); const reader = resp.body.getReader(); let buf = new Uint8Array(0); let complete = false; let anyAudio = false; for (;;) { const { done, value } = await reader.read(); if (done) break; const next = new Uint8Array(buf.length + value.length); next.set(buf); next.set(value, buf.length); buf = next; for (;;) { if (buf.length < 4) break; const length = new DataView(buf.buffer, buf.byteOffset, 4).getUint32(0); if (length === 0) { // end-of-stream marker complete = true; buf = buf.subarray(4); continue; } if (buf.length < 4 + length) break; const frame = buf.slice(4, 4 + length); buf = buf.subarray(4 + length); if (hasSamples(frame)) { anyAudio = true; yield frame; } // skip heartbeats } } if (!complete) throw new Error('stream truncated — discard this audio'); // Kết thúc sạch nhưng không một mẫu nào: chúng tôi cũng không tính tiền ca // này. Coi nó là thành công là ghi sổ lệch với hoá đơn. if (!anyAudio) throw new Error('stream carried no audio — not billed, do not save'); } ``` With the default WAV framing, each yielded frame is a complete file — in a browser you can feed them straight to `decodeAudioData` and queue the results. The `hasSamples` guard above is what makes that safe: without it a heartbeat decodes to a zero-length buffer and quietly joins the queue. ## OpenAI-style streaming If you are driving this from an OpenAI client, skip the framing entirely: ```python with client.audio.speech.with_streaming_response.create( model="tts-1", voice="Ngọc Lan", input="…", response_format="pcm", extra_body={"stream_format": "audio"}, ) as resp: resp.stream_to_file("speech.pcm") # raw s16le, 24 kHz ``` `stream_format: "audio"` gives you the audio bytes as chunked transfer encoding — append them and you have the whole thing. `stream_format: "sse"` gives Server-Sent Events: `speech.audio.delta` events carrying base64 audio, ending in exactly one `speech.audio.done` or `speech.audio.error`. As with the native stream, **`speech.audio.done` is the proof the audio is whole**; if it never arrives, discard what you have — you were not billed. `stream_format` accepts `pcm` and `ulaw` only, for the reason in the warning above: the other formats cannot be concatenated into a playable result. For a complete mp3 or opus file, make an ordinary non-streaming request. ## Latency, honestly **On the standard plans** — Starter through Hội viên, everything that shares the pooled lane — first audio lands in roughly **1–2 seconds** end to end. That is fast enough for read-aloud, dubbing, IVR prompts and assistants that tolerate a beat before speaking. It is **not** the 200–300 ms class that hard real-time conversational agents expect, and no amount of client tuning moves it: the wait is the queue plus the opener chunk, not your connection. **Enterprise is the exception, by contract rather than by luck.** That tier runs on a dedicated GPU node with a lane nobody else shares, so the latency ceiling becomes something we agree on — down to **≤ 250 ms** — instead of whatever the pool happens to be doing. It is a ceiling written into the deal and sized to your traffic, not a number you can read off the shared endpoint and expect to hold. If you are building a real-time agent, talk to us before you design around the 1–2 second figure above. --- # Errors Source: https://docs.vieneu.io/docs/cloud-api/errors Statuses tell you how bad it was. Most refusals additionally carry a `code`, and where one exists it is the only part of the body safe to branch on — messages are prose, some of it Vietnamese, and all of it subject to rewording. The important thing to know first is that **`code` is not universal**. Most of the ways a `/v1` call can be refused carry one; a few give you a status and a sentence. This page says which is which, because a client written on the assumption that every error has a `code` will silently fall through to its default branch on the ones that do not. ## Two body shapes Native endpoints return the platform's shape: ```json { "statusCode": 403, "message": "Insufficient tokens", "traceId": "0f7c…" } ``` `POST /v1/audio/speech` returns OpenAI's shape instead, so an OpenAI client can parse it without an adapter: ```json { "error": { "message": "…", "type": "rate_limit_exceeded", "param": null, "code": null } } ``` That is the whole difference: one route, one envelope. Everything else on `/v1` — including `/v1/vapi/speech`, which serves a third party — uses the first shape. Every native body repeats the status as `statusCode` and adds `code` and its extras where the refusal carries one: ```json { "statusCode": 429, "message": "Server is at capacity. Please try again in a few minutes.", "code": "QUEUE_FULL", "traceId": "0f7c…" } ``` A body that carried a `code` used to arrive without `statusCode` — so gaining a machine-readable reason cost you a field you had always been able to read. That gap is closed: the field is on every native body now. The HTTP status line is still the authoritative copy — read it from the response where you can. ## Refusals that carry a `code` ### Quota and grant The six routes that bill per character or per second — `POST /v1/tts`, `/v1/tts/stream`, `/v1/dialogue`, `/v1/dub`, `/v1/srt` and `/v1/vapi/speech` — answer a quota refusal with a code, and the two that recover on their own say when: | `code` | Status | What it means | What to do | |---|---|---|---| | `GRANT_EXPIRED` | 403 | The token package behind the key has passed its expiry date. | Renew. **Retrying will not help**, and neither will topping up. | | `GRANT_TOKENS_EXHAUSTED` | 403 | The grant's balance is smaller than this request's price. Carries `remaining` and `required`. | Top up. Retrying will not help. | | `GRANT_DAILY_LIMIT` | 429 | The plan's daily cap is spent. Carries `resetAt`, `remaining`, `required`. | Sleep until `resetAt`. | | `GRANT_WEEKLY_LIMIT` | 429 | The plan's weekly cap is spent. Carries `resetAt`, `remaining`, `required`. | Sleep until `resetAt`. | | `CONCURRENT_CONFLICT` | 429 | Two of your own requests raced for the same balance and this one lost. Carries `retryAfterSeconds: 1`, also sent as `Retry-After: 1`. | Wait the second it names, then retry — this one clears. | `resetAt` is ISO 8601, and it replaces parsing the timestamp out of the message. Take it over a guessed backoff, and over an "upgrade your plan" prompt for something that resets in an hour. The 403/429 split is the coarse version of the same decision: **403 means retrying will not help** even though it looks transient, and 429 means it will, eventually. What the status cannot always tell you is *when*. The daily and weekly caps carry **no `Retry-After`** — they carry `resetAt` instead, which is the honest answer for a wait measured in hours. The refusals that clear in seconds do carry it: `CONCURRENT_CONFLICT` (`Retry-After: 1`) and the capacity codes below (`Retry-After: 15`). `code` is what separates a race you retry from a cap you wait out, and `resetAt` is what says how long the wait is. `FREE_DAILY_LIMIT` and `GRANT_INACTIVE` appear in the platform's internal code list but not here: an API key always bills a grant, so the free tier is unreachable, and a grant that is not active answers 403 with prose and no code. ### Capacity | `code` | Status | What it means | What to do | |---|---|---|---| | `QUEUE_FULL` | 429 | The asynchronous job queue is at its depth limit. Carries `retryAfterSeconds` (`Retry-After`). | Wait what the header says and resubmit — with the same `Idempotency-Key` if you sent one. Nothing was charged. | | `USER_QUEUE_FULL` | 429 | *Your* account already holds its ceiling of queued jobs (30 by default). Carries `retryAfterSeconds`. | Poll what you have queued, then resubmit. Nothing was charged. | | `STREAM_CONCURRENCY` | 429 | You already hold the maximum number of open streams. | Close one, or wait for one to finish. Nothing was charged. | | `STREAM_BUSY` | 503 | Every node is healthy but streaming capacity is used up. | Carries `"fallback": "generate"`, and that is a real instruction — the queued path has separate capacity. See [How a stream ends](./streaming.md#how-a-stream-ends). | `QUEUE_FULL` and `USER_QUEUE_FULL` reach you from `POST /v1/tts` only; the synchronous routes do not queue. ### Idempotency Only on `POST /v1/tts`, and only when you sent an `Idempotency-Key` — see [Retrying safely](./overview.md#retrying-safely). | `code` | Status | What it means | What to do | |---|---|---|---| | `IDEMPOTENCY_KEY_INVALID` | 400 | The header is not 8–128 printable ASCII characters without whitespace. | Send a UUID. | | `IDEMPOTENCY_IN_FLIGHT` | 409 | A request with this key is still being processed. Carries `retryAfterSeconds: 1` (`Retry-After: 1`). | Retry after the second, with the **same** key. A new key can charge you twice. | | `IDEMPOTENCY_KEY_REUSED` | 422 | This key was already used with a different request body. | Use a fresh key for a different request; keys are per request, not per session. | ### Cloning Cloning moved to the web Studio. The enrolment routes — `POST /v1/voices`, `DELETE /v1/voices/{voiceId}`, `POST /v1/clone`, `POST /v1/upload` and `POST /v1/prepare` — answer one refusal now, before any billing or GPU work: | `code` | Status | What it means | What to do | |---|---|---|---| | `CLONE_WEB_ONLY` | 410 | Cloned voices are created at [vieneu.io/#/clone](https://vieneu.io/#/clone), not through the API. Carries `cloneUrl`. | Make the voice in the Studio once. Nothing else in your integration changes: it is then listed by `GET /v1/voices` with your key and its `clone_…` id works as `voiceId` on `POST /v1/tts` and `/v1/tts/stream`. | 410 rather than 404 on purpose: the endpoints existed and were documented, so the status says the surface moved rather than that you mistyped a path. The older clone codes (`CLONE_QUOTA_EXCEEDED`, `CLONE_MONTHLY_CAP`, `CLONE_DAILY_CAP`, `CLONE_REF_*`, `CLONE_TRANSCRIPT_DENSITY`, `CLONE_TOKENS_INSUFFICIENT` and the rest) belong to those routes and can no longer reach an API key — the same limits still apply to the Studio, where they are shown as you clone. Generating **with** a cloned voice is not part of this and carries no clone codes. A `clone_…` id that is missing or belongs to someone else fails as an ordinary `400` with a message and no `code` — see [Cloned voices](./streaming.md#cloned-voices). ## Refusals that do **not** carry a code ### The OpenAI-compatible route `POST /v1/audio/speech` carries none of the codes above, despite OpenAI's envelope having a field for it: its `code` is always `null` on this route. What you get instead is `type`, which is coarser on purpose. A quota refusal is typed from its status — `insufficient_quota` on the `403`s, which are the exhausted and expired grants, and `rate_limit_exceeded` on the `429`s, which are the spent caps and the races. So "out of credit" and "wait" survive the translation; `GRANT_EXPIRED` and `GRANT_TOKENS_EXHAUSTED` do not, and neither does the `resetAt` that would tell you how long to wait. Anything raised before the handler runs — a bad key, a field the validator rejects — is typed from the status alone, and that mapping folds `401` and `403` together as `authentication_error`. ### The public API never returns 402 There is no `402 Payment Required` anywhere on `/v1`. Every money path runs through the same deduction, and that step answers only `403` or `429`. The spec has carried a `402` example; the server does not send one. A client that treats `402` as "out of credit" will read the real `403` as a permissions bug and stop retrying for the wrong reason. ### Two text validations Before any tokens are spent, submitted text is checked for two things the model cannot usefully voice. Both come back as a plain `400` with a message and no `code`. | Message begins | Rule | |---|---| | `Text appears to be in an unsupported language…` | More than **34%** of the letters are outside the Latin script. Vietnamese diacritics and `đ` are Latin, so this fires on Chinese, Japanese, Korean, Cyrillic and the like — including an otherwise Vietnamese passage carrying a long non-Latin quotation. | | `Text does not look like readable words…` | Either one unbroken run of **more than 30 letters**, or **more than 60%** of the word-like tokens (four letters or longer, and at least three of them present) carry no vowel. | Punctuation breaks a letter run, so a URL, an email address or a file path measures as its parts and passes; a keyboard mash still measures as one run and does not. **These fire on three routes only:** `POST /v1/tts`, `POST /v1/dialogue` and `POST /v1/srt`. `POST /v1/tts/stream`, `POST /v1/audio/speech`, `POST /v1/vapi/speech` and `POST /v1/dub` do not run the check — so text one route rejects, another will synthesize and bill you for. Worth knowing before you treat a `400` from one route as proof the text is bad everywhere. ## Stream concurrency `POST /v1/tts/stream` refuses a new stream while you already hold too many: | Counted against | Ceiling | On exceeding | |---|---|---| | The token grant behind your API key | 4 | `429`, `code: STREAM_CONCURRENCY` | | A signed-in user in the web app | 2 | `429`, same code | The slot is taken **before** the balance is touched, so a refused stream leaves no mark on your tokens. Two details about the counting: it is per backend process, and the key is the **token grant**, not the API key — so several keys issued against one grant share the four rather than each getting four of their own. ## `X-Request-Id` Every response carries `X-Request-Id`. It is the id our logs are keyed by: one value spans the API request, the worker call it made, and every log line either produced. Most error bodies repeat the value as `traceId`, but **the header is the copy to read** — it is the only one that is always there. `traceId` is stamped by the exception filter, so a body written straight to the response never gets one: - **`POST /v1/audio/speech`**, on every error. The handler writes OpenAI's envelope itself, and the filter that would add `traceId` never runs — nor does OpenAI's shape have a field for it. - **`POST /v1/vapi/speech`**, on the `502` it sends when the worker produced no audio. That body is written mid-response and carries `statusCode` and `message` only. Quote it in a support request. Without it, "a request failed around 3pm" is a search; with it, it is a lookup. The value is always ours. The edge proxy sets `X-Request-ID` on every request it forwards, unconditionally, so an id you send on the way in is overwritten rather than adopted — log the one that comes back next to your own correlation id instead of expecting yours to survive. ## Status summary | Status | Meaning | |---|---| | 400 | Malformed request — an unknown voice, a bad format, text the validator refused, a clone rejection | | 401 | Missing, malformed or revoked API key | | 403 | Out of tokens, grant expired, or your plan does not include this engine or feature | | 413 | Uploaded file over 10 MB | | 409 | An `Idempotency-Key` whose first call is still running — retry with the same key | | 422 | Content refused by moderation (only when `aiRefine` is on), or an `Idempotency-Key` reused with a different body | | 429 | Rate limit, token quota, queue depth or stream concurrency | | 500 | Synthesis failed. The charge is refunded automatically | | 502 | Every worker for this engine failed. Refunded | | 503 | No worker for the requested engine or format, or streaming capacity is full — retry shortly | `429` is the one status with unrelated causes behind it, and the headers say which — see [Rate limits](./overview.md#rate-limits-and-request-ids). --- # Webhooks Source: https://docs.vieneu.io/docs/cloud-api/webhooks `POST /v1/tts` is asynchronous. Instead of polling `GET /v1/tts/{jobId}` until the status changes, register an HTTPS endpoint and we will POST a signed event to it the moment the job reaches a terminal state. Polling still works and is still the recovery path when a delivery fails — see [When we give up](#when-we-give-up). ## Events | Type | When | | --- | --- | | `tts.job.completed` | The job produced audio. The event carries a short-lived download URL. | | `tts.job.failed` | The job failed permanently after all retries. Tokens have been refunded. | Only `POST /v1/tts` produces jobs today, so those are the only two event types. Every other synthesis route (`/v1/audio/speech`, `/v1/tts/stream`, `/v1/dialogue`, `/v1/dub`, `/v1/srt`) answers inline and produces no event. ## Register an endpoint ```bash curl -X POST https://api.vieneu.io/api/v1/webhooks \ -H "Authorization: Bearer vn_sk_…" \ -H "Content-Type: application/json" \ -d '{ "url": "https://api.example.com/hooks/vieneu", "description": "Production receiver" }' ``` ```json { "id": "665f1a2b3c4d5e6f7a8b9c0d", "url": "https://api.example.com/hooks/vieneu", "events": ["tts.job.completed", "tts.job.failed"], "status": "ENABLED", "secretPrefix": "whsec_a1b2c3", "secret": "whsec_2f9c…" } ``` **`secret` is returned once and never again.** Store it now. If you lose it, call `POST /v1/webhooks/{id}/rotate-secret` and update your receiver. ### Destination requirements - **`https` only.** There is no `http` option, in any mode. Events carry a presigned download URL for your audio, which is a bearer credential — the signature protects integrity, not confidentiality. For local development use a tunnel (ngrok, Cloudflare Tunnel, the like) rather than an `http` URL. - **Public addresses only.** Private, loopback, link-local, CGNAT and reserved ranges are refused — when you register the endpoint, and again at the moment of every delivery. A hostname that resolves publicly at registration and privately later is refused at connect time. - **No redirects.** A `3xx` response is treated as a failed delivery. Point the endpoint at the final URL. - **No credentials in the URL.** `https://user:pass@…` is refused. ### Scoping to one API key Pass `apiKeyId` when registering to receive only the jobs submitted with that key. Omit it and the endpoint receives events for every key on the account. Useful for keeping a test key's traffic off your production receiver. ## The event body ```json { "id": "evt_9f2c1e043b8a4d218f770c1a2b3c4d5e", "type": "tts.job.completed", "api_version": "2026-09-01", "created": 1756684800, "attempt": 1, "data": { "job_id": "0f7c1a2b-3c4d-5e6f-7a8b-9c0d1e2f3a4b", "status": "completed", "voice_id": "Ngọc Lan", "engine": "v4", "character_count": 412, "token_cost": 1236, "duration_seconds": 27.4, "processing_time_ms": 8120, "audio_url": "https://…s3…?X-Amz-Signature=…", "audio_url_expires_at": "2026-09-01T11:00:00.000Z" } } ``` A failure looks like this: ```json { "id": "evt_1a4d…", "type": "tts.job.failed", "api_version": "2026-09-01", "created": 1756684800, "attempt": 1, "data": { "job_id": "0f7c1a2b-3c4d-5e6f-7a8b-9c0d1e2f3a4b", "status": "failed", "voice_id": "Ngọc Lan", "engine": "v4", "character_count": 412, "token_cost": 1236, "error": { "code": "synthesis_unavailable", "message": "No synthesis worker could complete the job. Tokens were refunded." }, "refunded": true } } ``` Branch on `error.code`, never on `error.message`. The codes are `synthesis_failed`, `synthesis_unavailable`, `synthesis_timeout`, `content_rejected` and `quota_exceeded`; the message is prose and may be reworded. ### About `audio_url` `audio_url` is minted fresh for **each delivery attempt** and expires one hour after that attempt — `audio_url_expires_at` tells you exactly when. It is deliberately shorter-lived than the 24-hour URL `GET /v1/tts/{jobId}` returns, because an event body may end up in your logs. Download promptly, or call `GET /v1/tts/{jobId}` for a fresh URL whenever you need one. The event carries everything you need to act on the job without a second API call. It does **not** carry the submitted text. ## Verifying the signature Every request carries a `Vieneu-Signature` header: ``` Vieneu-Signature: t=1756684800,v1=5257a869e7ecebeda32affa62cdca3fa51cad7e77a0e56ff536d0ce8e108d8bd ``` - `t` is the unix timestamp of **this delivery attempt**. - `v1` is `HMAC-SHA256(secret, "." + rawBody)`, hex-encoded. During a secret rotation the header carries **one `v1=` element per live secret**. Check the signature against each of them and accept if any matches. Four rules that matter: 1. **Use the raw request body bytes.** Not a re-serialised object. `JSON.parse` then `JSON.stringify` changes key order and unicode escaping, and the signature will fail intermittently in a way that is very hard to debug. Read the raw body before any JSON middleware touches it. 2. **Reject a stale `t`.** Use a tolerance of **300 seconds**. Do not use `0` — that disables the recency check entirely. 3. **Ignore any scheme that is not `v1`.** If a future element `v2=` appears, a verifier that accepts "any element that matches" can be downgraded. 4. **Compare in constant time.** `crypto.timingSafeEqual`, not `===`. ### Node.js ```js const crypto = require('crypto'); /** * Verify a Vieneu webhook signature. * * @param {Buffer|string} payload RAW request body — not a parsed object. * @param {string} header The `Vieneu-Signature` header value. * @param {string} secret Your `whsec_…` signing secret. * @param {number} toleranceSeconds Max age of `t`. 300 is the documented value. * @returns {boolean} */ function verifySignature(payload, header, secret, toleranceSeconds = 300) { const body = Buffer.isBuffer(payload) ? payload : Buffer.from(payload, 'utf8'); let timestamp = null; const signatures = []; for (const element of String(header || '').split(',')) { const idx = element.indexOf('='); if (idx === -1) continue; const key = element.slice(0, idx).trim(); const value = element.slice(idx + 1).trim(); if (key === 't') timestamp = Number(value); // Only v1. Ignoring unknown schemes is what stops a downgrade. else if (key === 'v1') signatures.push(value); } if (timestamp === null || !Number.isFinite(timestamp)) return false; if (signatures.length === 0) return false; // Replay window. A tolerance of 0 disables this check — don't. if (toleranceSeconds > 0) { const now = Math.floor(Date.now() / 1000); if (Math.abs(now - timestamp) > toleranceSeconds) return false; } const expected = crypto .createHmac('sha256', secret) .update(Buffer.concat([Buffer.from(timestamp + '.', 'utf8'), body])) .digest('hex'); return signatures.some((candidate) => { // timingSafeEqual throws on a length mismatch, so check length first — // the length of a hex digest is not itself a secret. if (candidate.length !== expected.length) return false; try { return crypto.timingSafeEqual( Buffer.from(candidate, 'hex'), Buffer.from(expected, 'hex'), ); } catch (err) { return false; } }); } ``` Wire it up in Express, taking care to keep the raw body: ```js const express = require('express'); const app = express(); app.post( '/hooks/vieneu', express.raw({ type: 'application/json' }), (req, res) => { const ok = verifySignature( req.body, // Buffer, thanks to express.raw req.get('Vieneu-Signature'), process.env.VIENEU_WEBHOOK_SECRET, ); if (!ok) return res.sendStatus(400); // Answer FIRST, work afterwards. We time out at 10 seconds. res.sendStatus(200); const event = JSON.parse(req.body.toString('utf8')); if (alreadyProcessed(event.id)) return; // at-least-once — dedupe on id void handle(event); }, ); ``` ### Python ```python import hashlib import hmac import time def verify_signature(payload: bytes, header: str, secret: str, tolerance_seconds: int = 300) -> bool: timestamp = None signatures = [] for element in (header or "").split(","): key, sep, value = element.partition("=") if not sep: continue key, value = key.strip(), value.strip() if key == "t": try: timestamp = int(value) except ValueError: return False elif key == "v1": # only v1; ignore any other scheme signatures.append(value) if timestamp is None or not signatures: return False if tolerance_seconds > 0 and abs(int(time.time()) - timestamp) > tolerance_seconds: return False expected = hmac.new( secret.encode("utf-8"), f"{timestamp}.".encode("utf-8") + payload, hashlib.sha256, ).hexdigest() return any(hmac.compare_digest(candidate, expected) for candidate in signatures) ``` ## Delivery guarantees **At-least-once, unordered.** Concretely: - **Duplicates are normal.** Design your receiver to be idempotent. Dedupe on `id`, which is stable for a given job and event type — the same terminal state re-emitted after a crash or a queue redelivery carries the same `id`. - **`t` and `v1` change between duplicates.** Every attempt is signed afresh. Never dedupe on the signature. - **Do not use `created` for ordering or deduplication.** It is a diagnostic timestamp. Two jobs' events can arrive in either order, and a retry of an older event can land after a newer one. - **`attempt` tells a retry from a first delivery.** It is 1-based, and the `Vieneu-Delivery` header identifies one specific HTTP attempt. ### Other headers | Header | Meaning | | --- | --- | | `Vieneu-Signature` | `t=…,v1=…` — see above. | | `Vieneu-Event-Id` | Same value as `id` in the body. Convenient for dedupe at the edge. | | `Vieneu-Event-Type` | Same value as `type` in the body. | | `Vieneu-Delivery` | Identifies ONE http attempt. Not a dedupe key. | ## What we expect from your endpoint - **Answer `2xx`.** Anything else — including any `3xx` — is a failed delivery. - **Answer within 10 seconds.** Acknowledge first, do the work after. - **Keep the response small.** We stop reading as soon as 64 KB has arrived (we finish the chunk in flight, so a little more may cross the wire) and discard it. We never log, store or return your response body. ## Retries and backoff A failed delivery is retried up to **6 attempts**, with the gaps growing each time. Measured against the queue library we actually run, they fire at: ``` t+0 t+0.5m t+2.0m t+5.5m t+13.0m t+28.5m ``` So the window is about **28.5 minutes** end to end, and the longest single gap between two attempts — the one before the last — is **15.5 minutes**. That second number is the one to build against. **Size any staleness or reconciliation threshold above 15.5 minutes.** A sweeper that treats a delivery as stranded after fifteen will keep finding deliveries that are simply waiting out their final backoff, and will re-enqueue work that was never lost. :::note This page used to say 15 minutes It described the schedule as "+30s, +1m, +2m, +4m, +8m, about 15 minutes". That was arithmetic on the wrong formula — the queue computes each delay as `(2^attempts − 1) × 30s`, not `30s × 2^attempts`, which roughly doubles both the window and every gap inside it. The numbers above were read off the installed library rather than recalled. ::: ## When we give up After the last attempt the event is **dead-lettered**: it will not be sent again. Two things happen: 1. The delivery row is marked `DEAD`, visible at `GET /v1/webhooks/{id}/deliveries`. 2. A notification appears in your Vieneu account. The audio is not lost. `GET /v1/tts/{jobId}` still returns the job and a fresh 24-hour download URL. If your receiver was down, reconcile from there. ## Self-diagnosis ```bash curl https://api.vieneu.io/api/v1/webhooks/{id}/deliveries \ -H "Authorization: Bearer vn_sk_…" ``` ```json [ { "id": "665f…", "eventId": "evt_9f2c…", "eventType": "tts.job.completed", "status": "DEAD", "attempts": 6, "lastStatusCode": 502, "lastError": "http_502", "lastDurationMs": 143, "lastAttemptAt": "2026-09-01T10:15:02Z", "deliveredAt": null } ] ``` `lastError` values you may see: | Code | Meaning | | --- | --- | | `http_4xx` / `http_5xx` | Your endpoint answered with that status. | | `redirect_refused` | Your endpoint returned a `3xx`. Point it at the final URL. | | `connect_timeout` / `response_timeout` | Your endpoint did not answer in time. | | `connection_refused` / `dns_failure` / `tls_failure` | We could not reach it. | | `blocked_private_address` | The URL resolves to a non-public address. | | `endpoint_unavailable` | The endpoint was deleted or disabled mid-flight. | | `audio_not_uploaded` | The audio had not finished uploading yet. Retried. | ## Rotating the signing secret ```bash curl -X POST https://api.vieneu.io/api/v1/webhooks/{id}/rotate-secret \ -H "Authorization: Bearer vn_sk_…" ``` The new secret is returned once. The previous secret keeps verifying for **24 hours**, and during that window each event carries one `v1=` per live secret — so you can deploy the new secret without dropping events signed with the old one. The verifier above already handles this: it accepts if any `v1` matches. --- # Vapi (voice agent) Source: https://docs.vieneu.io/docs/cloud-api/vapi [Vapi](https://vapi.ai) builds voice agents that answer and place phone calls. It speaks through a TTS provider of your choosing — and its `custom-voice` provider lets that be any HTTPS endpoint. VieNeu implements exactly what it expects, so your agent can answer in Vietnamese. This matters because Vietnamese is essentially absent from the realtime TTS vendors Vapi ships with. ## Configure the assistant ```json { "voice": { "provider": "custom-voice", "server": { "url": "https://api.vieneu.io/api/v1/vapi/speech?voiceId=Ngọc%20Lan", "secret": "vn_sk_your_key_here", "timeoutSeconds": 30 } } } ``` That is the whole integration. Two details are doing the work: **`secret` is your VieNeu API key.** Vapi sends it as the `X-VAPI-SECRET` header, and its config has no field for an `Authorization` header — so this is where the key goes. It is the same key, checked the same way, and it is billed to the same account. **The voice goes in the URL.** Vapi's request payload has no voice field, so pass `?voiceId=` (URL-encoded). List the options with `GET /v1/audio/voices`. Add `&engine=v4` for the premium engine if your plan includes it. Omit `voiceId` entirely and you get the engine's default voice. ## What happens on each utterance Vapi POSTs: ```json { "message": { "type": "voice-request", "text": "Xin chào, tôi có thể giúp gì cho bạn?", "sampleRate": 24000 } } ``` VieNeu answers `200` with `Content-Type: application/octet-stream` and raw mono 16-bit little-endian PCM at exactly that sample rate, streamed as it is generated rather than buffered — so the agent starts speaking sooner. Vapi asks for 8000, 16000, 22050 or 24000 Hz depending on the transport. All four are supported, and the response is always at the rate requested: raw PCM carries no rate of its own, so anything else would come out at the wrong pitch. ## Billing and behaviour Billed per submitted character, like the rest of `/v1`. Audio that gets cut short is refunded automatically. **AI refinement never runs on this endpoint**, regardless of your account settings. An agent's job is to answer quickly, and an extra model round-trip in front of every utterance is the opposite of that. Deterministic text preparation still runs, so Vietnamese is pronounced correctly. If your agent's text contains things you want read a particular way — currency, dates, product codes — normalize it in your prompt, where you can see the result. ## Latency, honestly **On the standard plans**, first audio leaves in roughly 1–2 seconds. That works for an agent that answers a question, reads a menu, confirms a booking. It is **slower than the 200–300 ms** the fastest English-only vendors reach, and it will be noticeable as a beat before the agent speaks. An Enterprise deal runs on a dedicated node where the latency ceiling is agreed per contract — down to ≤ 250 ms — which is the version of this you want if the beat is a dealbreaker. See [Latency, honestly](./streaming#latency-honestly). Two things worth doing: keep `timeoutSeconds` at 30 or above, and keep utterances short — a paragraph costs more waiting than three sentences. **Utterances are capped at 800 characters** (about 50 seconds of speech), and the cap exists for that reason: a longer turn cannot be delivered inside Vapi's 30-second deadline, so it would be cut off mid-word and billed anyway. If your agent produces longer replies, split them — Vapi will request each piece separately and the caller hears them back to back. If no worker starts answering within 12 seconds, the request fails with a `503` rather than holding the socket until Vapi gives up. That distinction matters: a status code your `fallbackPlan` can act on beats a timeout, which just looks like a provider that stopped responding. Set a `fallbackPlan` on the assistant if a missed utterance would be worse than a non-Vietnamese voice. ## Troubleshooting | Symptom | Cause | |---|---| | `401` | The `secret` is not a valid VieNeu API key, or the key was revoked. | | `400` with a voice message | `voiceId` is not in the catalogue for that engine — check `GET /v1/audio/voices`. | | `403` | The key's token grant is empty or expired, or your plan excludes the engine. | | `400` about `message.text` length | The turn was over 800 characters — split it. | | `503` naming a sample rate | No worker in the pool has been updated to encode raw PCM yet. We fail over across the fleet first, so this means all of them. Retry; contact us if it persists. | | `503` about no worker answering | Nothing started producing audio within 12 seconds. Usually a capacity spike; your `fallbackPlan` covers the turn. | | `502` about no audio | The worker accepted the request and then produced nothing. Not billed. | Every response carries `X-Request-Id`. Quote it and we can find the exact call. --- # Changelog Source: https://docs.vieneu.io/docs/cloud-api/changelog Changes to the Cloud API (`/api/v1`). Anything marked **⚠️ Behaviour change** can affect a running integration of yours even if you change nothing. --- ## 2026-10-02 — Long audiobooks: nothing lost, nothing stuck {#2026-10-02--audiobooks-resilience} ### ⚠️ Behaviour change **A chapter longer than the plan can render is refused when added.** A chapter is billed in one go, so on a plan with a daily or weekly token cap, a chapter costing more than the cap could never start: it waited for tokens forever. `POST /v1/audiobooks/{bookId}/chapters` now answers 400 `chapter_over_plan_limit` with `chapterCharsMax`, and every book reports `chapterCharsMax`. Plans without such a cap are unchanged (60,000 characters). ### New **`afterSegment` when continuing a chapter.** With `appendTo`, pass how many segments the chapter has; a different count answers 409 `segment_mismatch` with the words the chapter ends with, so a retried request cannot add its text twice and a lost one cannot leave a gap. Without it, a request repeating the chapter's last segments answers 409 `duplicate_segments`, and a new chapter identical to the last one 409 `duplicate_chapter`. The MCP tool `add_audiobook_chapter` takes `after_segment` and shows how each chapter now ends. See [Audiobooks](../integrations/mcp/audiobooks.md#limits). ### Changed **Failed chapters are made again on their own.** A chapter whose job fails — a deploy replacing the server twice during one long chapter, a voice server down — is refunded as before, then rendered again a few minutes later, at most twice, reusing the parts already made. A job that stops moving for 45 minutes is failed and refunded the same way, so it can no longer hold up every other book. Mastering to MP3 retries with growing waits for about an hour before handing over the WAV, and a take cut short in transit, or an MP3 shorter than its take, is refused instead of delivered. **MP3 peaks stay under −3 dBTP.** Mastering now aims half a decibel lower and adds a peak limiter before the MP3 encoder, which lifts peaks a little: two chapters of a test book rendered on production had reached −2.8 dBTP. Loudness is unchanged at −19 LUFS. --- ## 2026-10-01 — Audiobooks, over the API and in the MCP server {#2026-10-01--audiobooks} ### New **Audiobooks.** A book is a narrator and up to 30 character voices reading chapters written as a script — narration, and each line of dialogue with the character who says it. `POST /v1/audiobooks` creates one, `POST /v1/audiobooks/{bookId}/chapters` adds chapters, `POST /v1/audiobooks/{bookId}/start` queues it and `GET /v1/audiobooks/{bookId}` follows it, with a download link for every finished chapter. Books are made in the background, one chapter at a time, while the GPU fleet is idle, so they take hours rather than seconds. Each chapter is billed per character when it starts, refunded if it fails, mastered to MP3 (−19 LUFS, peaks ≤ −3 dBTP, silence at both ends, ID3 tags) and saved to the Library. LIVE keys and catalogue voices only. See [Every operation](./overview.md#every-operation). **Audiobooks in the MCP server.** Six tools and a chapter player let Claude or ChatGPT turn a text the user provides into a multi-voice audiobook. See [Audiobooks](../integrations/mcp/audiobooks.md). **"What can VieNeu do?" in the MCP server.** A `list_capabilities` tool (a menu card in Claude and ChatGPT: click an example to send it) and five ready-made prompts. Optional tool arguments now also accept `null`. See [MCP server](../integrations/mcp/index.md#prompts). **Connected before this release? Reconnect.** Apps keep the tool list they fetched when they connected, so the new tools appear after you disconnect VieNeu and connect again. From now on the server's `serverInfo.version` changes whenever its tools do (`1.2.0+`), for apps that refresh on it. See [Troubleshooting](../integrations/mcp/troubleshooting.md). **Voice search that reads a casting brief.** In the MCP server, `list_voices` takes `search` the way assistants write it, "giọng nam trầm ấm miền Bắc": gender and region come from the voice's fields, a word written with diacritics must match them, words such as "giọng" are ignored, and when no voice has every word the closest come back with what each lacks. See [MCP server](../integrations/mcp/index.md#list_voices). --- ## 2026-10-01 — MP3 and an inline player in the MCP server; `X-Vieneu-Job-Id` on `/v1/audio/speech` {#2026-10-01--mcp-mp3-and-inline-player} ### New **The MCP server returns MP3 and plays it in the chat.** Texts up to 1,500 characters now come back as MP3 (about a tenth of the WAV size) instead of WAV, and in Claude the result shows an audio player right in the conversation. See [MCP server](../integrations/mcp/index.md#the-inline-player). **`POST /v1/audio/speech` names its saved copy.** A successful non-streaming response now carries `X-Vieneu-Job-Id`: the id under which VieNeu keeps a copy of the audio you just received. A moment later `GET /v1/tts/{jobId}` answers `completed` with a presigned `audioUrl` for the same bytes — useful when you want a shareable link rather than the raw file. The header is exposed to browsers through CORS. **`GET /v1/balance`.** How many tokens the calling key can spend right now, the plan it bills, and that plan's daily and weekly caps with their reset times. Not billed. The MCP server exposes the same thing as the `get_token_balance` tool and adds the remaining balance to every finished synthesis. **MCP sign-in works with ChatGPT.** The authorization server now accepts a client metadata document that prefers `private_key_jwt` as long as it also allows `none` (ChatGPT's does), and dynamic registration accepts `client_secret_post` / `client_secret_basic`. --- ## 2026-09-30 — Hosted MCP server: VieNeu inside Claude and ChatGPT {#2026-09-30--hosted-mcp-server} ### New **`https://api.vieneu.io/mcp`** — add this URL to Claude, ChatGPT, Cursor or Claude Code, sign in with your VieNeu account, and ask for Vietnamese speech in plain language. The assistant searches voices, synthesizes and returns a download link. No package to install and no API key to paste: the assistant signs in through OAuth 2.1 and you approve it on a VieNeu page. Tool calls are billed exactly like the equivalent `/v1` calls. See [MCP server](../integrations/mcp/index.md) and the per-app guides for [Claude](../integrations/mcp/claude-ai.md), [ChatGPT](../integrations/mcp/chatgpt.md) and others. Connected apps are listed, and can be disconnected, on the Developer page. --- ## 2026-09-25 — A cache hit on `POST /v1/tts` answers `completed` ### New **`POST /v1/tts` returns the audio URL at once when the audio already exists.** Identical requests — the same text, voice, speed, emotion and engine, from any account — are served from a content cache rather than synthesized again. Such a call used to answer `status: "queued"` like every other, sending you off to poll for audio that was already there: one full poll interval spent on nothing. It now answers `status: "completed"` with `audioUrl`, `audioUrlExpiresIn` and `duration` — exactly the body `GET /v1/tts/{jobId}` would give — so there is nothing to poll. Billing is unchanged (a cache hit is charged, as before). If your client always polls after a submit, nothing breaks: the poll returns the same `completed` body. To take the shortcut, branch on `status` in the submit response: ```js const job = await submit(text, voiceId); const done = job.status === 'completed' ? job : await pollUntilDone(job.jobId); ``` A submit that is *not* a cache hit still answers `queued` — and so does an `Idempotency-Key` replay of an earlier submit, whatever that job's state; poll it. The cached row's bytes are occasionally still on their way to storage, in which case the response stays `queued` and the poll reports `completed` a moment later. --- ## 2026-09-24 — `v3` retired from the cloud API ### ⚠️ Behaviour change **The `v3` engine is gone from `/api/v1`.** From 2026-09-24 08:00 (GMT+7) the cloud API renders on `v4` only. Any request that names `"engine": "v3"` — on `POST /v1/tts`, `/v1/tts/stream`, `/v1/audio/speech`, `/v1/dialogue`, `/v1/dub`, `/v1/srt` and their multipart variants — is refused with `400` before anything is billed, and the message says so: ``` Engine "v3" was retired on 2026-09-24. Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues share no ids. ``` A request that names no engine was already landing on `v4`, the default, and is not affected — unless it carries a `v3` voice id, which is the next point. **What the catalogue endpoints say now.** `GET /v1/engines` lists a single engine, `{ "key": "v4", "isDefault": true, "sampleRate": 48000, "billingMultiplier": 3, "features": ["generate", "clone", "dialogue", "stream", "dub", "srt"] }`, with `"globalMultiplier": 1.3` beside it. `GET /v1/voices` lists `v4` voices only, plus your own `v4` clones when you send a key; `v3` voices, and clones enrolled on `v3`, are no longer listed. `GET /v1/voices?engine=v3` is a `400` with the message above. The retired catalogue's ids — mostly opaque `vieneu-…` slugs — shared almost no space with `v4`'s display names, so a hard-coded `v3` voice keeps failing on `voiceId` even once you drop `engine`: take a fresh id from `GET /v1/voices?engine=v4`. **On the OpenAI-compatible route**, `"model": "vieneu-v3"` answers `400` with `Model 'vieneu-v3' is retired. Engine "v3" was retired on 2026-09-24. …`. `tts-1`, `tts-1-hd`, `gpt-4o-mini-tts`, `vieneu` and `vieneu-v4` all resolve to `v4`. The `emotion` field (`natural` / `storytelling`) was `v3`-only; it is still accepted, and ignored — `v4` has no reading styles. **`GET /v1/emotion-tags` with no `engine`** now answers for the default engine, `v4`: no styles and three inline cues, `[cười]`, `[thở dài]` and `[hắng giọng]`. It used to answer with `v3`'s set. `?engine=v3` is a `400`. The `/v1/audio/…` equivalents behave the same way. **Clones enrolled on `v3`** — anything created before 2026-08-28, when enrolment moved to `v4` — can no longer be rendered through the API, because the only engine that could read their reference clip is the one that now answers `400`. Naming one as `voiceId` (or as the `voice` of `/v1/audio/speech`) answers `400` with `code: "CLONE_ENGINE_RETIRED"` and the voice's `engine`, instead of quietly rendering the clip on `v4` in a voice you never enrolled. Re-enrol the voice in the [Studio](https://vieneu.io/#/clone); the new `clone_…` id works on `POST /v1/tts` and `POST /v1/tts/stream` exactly as before. **What is not affected.** Audio already generated on `v3` is untouched — library entries and download links keep working. And this is a cloud-API change only: the on-device **v3 Turbo** in the Python SDK and the Windows app runs on your own machine and is unchanged. --- ## 2026-09-23 — Cloning moved to the web Studio ### ⚠️ Behaviour change **The clone-enrolment endpoints are closed.** `POST /v1/voices`, `DELETE /v1/voices/{voiceId}`, `POST /v1/clone`, `POST /v1/upload` and `POST /v1/prepare` now answer `410` with `code: "CLONE_WEB_ONLY"` and a `cloneUrl`. Nothing is billed and no GPU work starts — the refusal happens before either. Create the voice at [vieneu.io/#/clone](https://vieneu.io/#/clone) instead. Why the web: enrolling well is a guided job. The Studio denoises the clip, transcribes it with Whisper and lets you hear the voice before it is saved. A reference that is noisy, mistranscribed or the wrong length degrades **every** later generation with that voice, not just the call that created it — and an API caller had no way to see that before spending the enrolment fee. **Nothing changes for generation.** A voice cloned on the web is listed by `GET /v1/voices` when you send your key (`"kind": "cloned"`) and its `clone_…` id remains an ordinary `voiceId` on `POST /v1/tts` and `POST /v1/tts/stream`, at the same price as a catalogue voice. Integrations that only *use* clones need no change at all. The clone routes are gone from [`/openapi.json`](pathname:///openapi.json) and the [API reference](/api-reference), so generated clients will drop them on their next regeneration. The old clone error codes (`CLONE_MONTHLY_CAP`, `CLONE_QUOTA_EXCEEDED`, `CLONE_REF_*`, …) belonged to those routes and can no longer reach an API key; the same plan limits still apply in the Studio. See [Errors → Cloning](./errors.md#cloning). --- ## 2026-09-16 — Safe retries and batch polling on `POST /v1/tts` ### New **`Idempotency-Key` on `POST /v1/tts`.** The server creates the job and charges the tokens *before* it answers, so a client whose request timed out — or was cut by a `502` from the edge during a deploy — could never tell "never happened" from "happened, answer lost". Send a client-chosen key (8–128 printable ASCII characters; a UUID is ideal) and a retry with the same key and the same body returns the **first** job: same `jobId`, no second charge, and the response header `Idempotent-Replayed: true`. Keys are scoped to your API key and kept for 24 hours. Same key with a different body is refused with `422 IDEMPOTENCY_KEY_REUSED`; the same key while the first call is still running is `409 IDEMPOTENCY_IN_FLIGHT` with `Retry-After: 1` — retry with the **same** key, never a new one. Requests without the header behave exactly as before. See [Retrying safely](./overview.md#retrying-safely). **`GET /v1/tts?ids=a,b,c`** — poll up to 50 jobs with one request. Each entry has exactly the shape of `GET /v1/tts/{jobId}`; ids that are unknown or not yours are listed in `missing` rather than failing the batch. Not throttled, like the single poll. Prefer it whenever more than one job is in flight: the edge proxy caps concurrent connections per source address, and one batch call costs the server one database read where eight single polls cost sixteen. **`retryAfterSeconds` in error bodies, mirrored as `Retry-After`.** `CONCURRENT_CONFLICT` now says how long to wait (one second) instead of leaving you to guess; `QUEUE_FULL` and `USER_QUEUE_FULL` already carried it, and it is now documented on the error body. `USER_QUEUE_FULL` — the per-account ceiling on queued jobs — is also new to this document, though not to the API. ### Behaviour change — additive **A full queue no longer touches your balance.** `POST /v1/tts` used to charge, discover the queue was full, and refund inline. It now asks the queue first, so a `QUEUE_FULL` / `USER_QUEUE_FULL` answer moves no money — which is also what its message has always promised. --- ## 2026-09-13 — `GET /v1/usage` ### New **`GET /v1/usage`** — your own usage, from the same per-call ledger the platform keeps: calls by outcome, tokens, characters, seconds of audio, generation time (average, p95), broken down per day, per route, per voice and per error code. `scope=key` (default) is the calling key; `scope=account` merges every key on the account. Window up to 92 days. See [Seeing what you were billed for](./overview.md#seeing-what-you-were-billed-for). Nothing else changed for callers. Behind it, every `/v1` call is now recorded with its key, route, voice, cost and outcome — the reason a usage question can be answered exactly rather than estimated. --- ## 2026-09 — Error bodies, quota codes, and the clone engine gate ### ⚠️ Behaviour change — additive, nothing removed Read this line first: **no field was removed, renamed, or changed in type or meaning.** All three changes below only put more into a body that was already being sent. The only integration that can notice is one that rejects unknown fields — a strict schema, a Go struct decoded with unknown-field checking — and that is the reason this is filed as a behaviour change rather than under *New*. **Every native `/v1` error body now carries `statusCode`.** It used to appear only on the bodies that had nothing else in them: the moment a refusal gained a machine-readable `code`, it lost `statusCode` in the same move, so gaining a reason cost you a field you had always been able to read. One shape now, on every error the platform envelope covers. The HTTP status line remains the authoritative copy. `POST /v1/audio/speech` is unaffected — it answers in OpenAI's envelope, which has no `statusCode`. **Quota refusals now carry `code`, plus `resetAt`, `remaining` and `required` where the deduction supplies them.** Eight refusal sites across the seven billed routes — `/v1/tts`, `/v1/tts/stream`, `/v1/dialogue`, `/v1/dub` (two of them: the affordability gate before any GPU work, and the charge after it), `/v1/clone`, `/v1/srt` and `/v1/vapi/speech` — used to throw a bare sentence, so the only way to tell "top up" from "wait an hour" was to match English prose. They now carry `GRANT_EXPIRED`, `GRANT_TOKENS_EXHAUSTED`, `GRANT_DAILY_LIMIT`, `GRANT_WEEKLY_LIMIT` or `CONCURRENT_CONFLICT`, and the two limits that lift on their own carry an ISO 8601 `resetAt` saying when. `message` is unchanged, so a client reading only that keeps working. See [Errors](./errors.md#quota-and-grant) for what each code means and what to do about it. **`POST /v1/voices` now gates the engine that actually runs.** Enrolment has been `v4`-only since 2026-08-28, but the entry check was still asserting things about `v3` — so an enrolment that named no engine could be refused because of the state of an engine it was never going to use. The check now runs against the engine the request will really enrol on, and only when you named one explicitly. - *What you will see:* an enrolment that omits `engine` — the normal path — no longer fails for reasons that have nothing to do with it. - *What has not changed:* an explicit `"engine": "v3"` is still refused with `400 CLONE_ENGINE_NOT_ALLOWED`, and an explicit engine your plan does not cover is still `403`. --- ## 2026-09 — Webhooks ### New **Webhooks for asynchronous work.** Register an `https` endpoint with `POST /v1/webhooks`, and instead of polling `GET /v1/tts/{jobId}` you receive `tts.job.completed` or `tts.job.failed` the moment the job ends. The event carries everything you need to act on it without a second call, plus a download URL that lives for one hour. - Signed with HMAC-SHA256 in a `Vieneu-Signature: t=…,v1=…` header — the same shape Stripe uses, so the signature verifier you have already written carries over almost untouched. A copy-paste-and-run `verifySignature` is in [Webhooks](./webhooks.md). - **At-least-once and unordered delivery.** Duplicates are normal: dedupe on `id` (stable per job and event type). Do not dedupe on the signature — every attempt is signed afresh, so `t` and `v1` both change. - A failed delivery is retried up to 6 attempts, at `t+0`, `+0.5m`, `+2.0m`, `+5.5m`, `+13.0m` and `+28.5m` — about **28.5 minutes** for the whole run, and the longest single gap between two attempts is **15.5 minutes**. (This entry originally said "about 15 minutes"; that was arithmetic on the wrong backoff formula, see [Webhooks](./webhooks.md#retries-and-backoff).) Size any reconciliation threshold **above 15.5 minutes**. Once the attempts run out the event is **dead-lettered**: it will not be sent again, a notification appears in your account, and `GET /v1/webhooks/{id}/deliveries` tells you why. The audio is **not** lost — `GET /v1/tts/{jobId}` still returns the job and a fresh download URL. - The endpoint must be `https` and must point at a public address. Private ranges, loopback, link-local, CGNAT and reserved ranges are all refused — both at registration and again at the moment of delivery. Any `3xx` response counts as a failed delivery. - Rotate the signing secret with `POST /v1/webhooks/{id}/rotate-secret`. The old secret keeps working for 24 hours and each event carries one `v1=` per live secret, so you can roll out the new one without dropping a single event. Nothing changes behaviour here: polling works exactly as before, and an account that has registered no endpoint sees no difference at all. --- ## 2026-08 — Output formats, OpenAI compatibility, and Vapi ### New **Cloned voices work on `POST /v1/tts/stream`.** Pass a `clone_…` id (from `GET /v1/voices`, created with `POST /v1/voices`) as `voiceId` and it streams, just like a built-in voice. The streaming path used to refuse every `clone_…` id while the web app had been streaming cloned voices for a long time — the worker supported it all along; only the public surface blocked it. - Billing is **unchanged**: still submitted characters × the engine multiplier. Clones carry no surcharge of their own; the only charge is the one-off enrolment fee on `POST /v1/voices`. - A cloned voice is bound to the engine it enrolled on, so the "streaming is `v4` only" rule still applies. New clones all enrol on `v4`. - This is not a behaviour change: whatever worked yesterday still works exactly the same. **Multiple output formats.** `POST /v1/tts/stream` and `POST /v1/audio/speech` now take `mp3`, `opus`, `pcm` (raw, headerless) and `ulaw` (G.711 8 kHz for telephony), with a choice of sample rate. Before, WAV at 48 kHz was the only option. mp3 is roughly 6× smaller than WAV. **`POST /v1/audio/speech` is a real drop-in for OpenAI.** Point an OpenAI SDK at `https://api.vieneu.io/api/v1` with your API key and it works, with nothing else to change. `speed` is now applied, `model` selects the engine, and `stream_format` lets you take audio while it is still being generated (chunked, or OpenAI-style SSE). `GET /v1/audio/voices` was added so OpenAI clients can list voices. **`POST /v1/vapi/speech`** — the custom-voice webhook for [Vapi](https://vapi.ai). Point your assistant's `custom-voice` at this URL and your voice agent speaks Vietnamese. Details in [Vapi integration](./vapi.md). **Rate-limit headers on every response.** `X-RateLimit-Limit`, `X-RateLimit-Remaining`, `X-RateLimit-Reset`, and `X-Request-Id` — quote `X-Request-Id` when you contact support, it is the id our logs are keyed by. **Public documentation.** The Cloud API section on docs.vieneu.io, with reference decoders in Python and JavaScript for streaming. ### ⚠️ Behaviour change **`aiRefine` now defaults to OFF, not ON.** The AI refinement step (content moderation + pronunciation normalisation for formulas, acronyms and mixed-in English) used to run by default on `/v1` and carried a surcharge. You now have to ask for it explicitly with `"aiRefine": true`. - *What you will see:* cost per request **down**, latency **down**, and your text read **exactly as you sent it**. - *If you want the old behaviour:* add `"aiRefine": true` to the request. On `POST /v1/srt` (multipart) it is an `aiRefine=true` field. - Deterministic text preparation (chemistry readings, UPPERCASE folding, the sea-g2p phoneme layer) runs in every case — Vietnamese is still pronounced correctly. **The `response_format` default on `/v1/audio/speech` moves from `wav` to `mp3`.** That is OpenAI's default, and an OpenAI client that leaves the field empty is expecting mp3. If you need WAV, send `"response_format": "wav"`. **Rate limits change unit: from per IP to per API key.** They used to count IP addresses, which meant an office sharing one network had to split a limit between them, while one key spread across several machines multiplied it. They now count the key — the thing you actually bought. Alongside that, the ceiling on the three main synthesis routes was **raised to 300 requests per minute** so that nobody ends up with less than before. This is the **application-tier** limit. In front of it sits an edge proxy that still counts source addresses, and that tier has not changed. How to tell them apart: a `429` carrying `X-RateLimit-*` is your key hitting the application ceiling; a bare `429`/`503` with no headers at all is the edge proxy blocking by IP — shared with everyone else behind that address. See [Rate limits](./overview.md#rate-limits-and-request-ids). ### Fixed - Streaming through `api.vieneu.io` is no longer buffered by the proxy — audio arrives as it is generated, the way it was designed to. - `POST /v1/tts/stream` now reports the real format and sample rate in the `X-Output-Format` / `X-Sample-Rate` headers, instead of echoing back what you asked for. - Truncated streams are refunded more reliably: a stream that ends correctly but carries **no audio at all** is now refunded too. ### Known limitations `stream_format` (and streaming in general) should only be used with `pcm` and `ulaw`. Audio frames are encoded independently of one another, so concatenated mp3 has a gap of silence at every join, and opus becomes a chain of Ogg streams that most browsers only play the first of. If you need a complete mp3 or opus file, send an ordinary, non-streaming request. --- # Integrations overview Source: https://docs.vieneu.io/docs/integrations/overview :::info The n8n node is not published yet The n8n node (`n8n-nodes-vieneu` 0.1.0) is marked `private` and has never been published to npm. It is built and working, but the repository it lives in is not public, so there is no install command you can run today — **contact VieNeu to get the package**, and watch the [changelog](../cloud-api/changelog) for the release. Nothing else here depends on it. The MCP server, the OpenAI-compatible endpoint, the raw API, webhooks and Vapi are live and need nothing installed. ::: This page routes you to the right one. Each destination owns its own detail; this one deliberately repeats none of it. ## Which path is yours | What you already have | How VieNeu plugs in | Where | |---|---|---| | An app that speaks the **OpenAI TTS protocol** — Open WebUI, SillyTavern, LobeChat, LiteLLM, anything on the OpenAI SDK | Three settings: base URL, key, voice id. No plugin, no adapter. | [OpenAI-compatible apps](./openai-clients.md) — per-client setup and troubleshooting by symptom · [endpoint reference](../cloud-api/openai-compatible) | | An **AI assistant** — Claude, ChatGPT, Cursor, Claude Code | A hosted MCP server: add `https://api.vieneu.io/mcp`, sign in with your VieNeu account, ask for speech in plain language. Nothing to install. | [MCP server](./mcp/index.md) | | **n8n** | A community node (self-hosted n8n, built from source), or two ready-made workflows built from n8n's own HTTP Request node that install nothing and run on n8n Cloud. | [n8n node](./n8n.md) | | A **voice agent or phone system** | Vapi has a dedicated webhook route. LiveKit Agents, Pipecat and anything else with an OpenAI TTS plugin use `/v1/audio/speech` with `response_format: "pcm"`. | [Vapi](../cloud-api/vapi) · [LiveKit / Pipecat](./openai-clients.md#livekit-agents) · [Streaming](../cloud-api/streaming) | | **Your own code** | Call `/api/v1` directly — that surface is much larger than the OpenAI-shaped one. | [Cloud API overview](../cloud-api/overview) · [API reference](/api-reference) | | Something that must **react when a job finishes** | Register a webhook instead of polling. | [Webhooks](../cloud-api/webhooks) | ## What every path shares Base URL `https://api.vieneu.io/api/v1`. One key, under either header name: ``` Authorization: Bearer vn_sk_... X-API-Key: vn_sk_... ``` `vn_sk_` keys are live; `vn_test_` keys behave identically but cap each request at 100 words. Any other string — `sk-…`, `none`, an empty placeholder — is a **401 before any lookup**, so a client that insists on a non-empty key field must be given a real one. A `Bearer` token containing a dot is read as a JWT and ignored, which is why pasting one produces "API key required" rather than "invalid key". A third header, `X-VAPI-SECRET`, is accepted on every `/v1` route — last in precedence, so an explicit header always wins. It exists because a Vapi assistant's configuration can set neither of the other two. See [Authentication](../cloud-api/overview#authentication). Synthesis is billed per **submitted character**, minimum 50, times the engine's multiplier; a request that produces no audio is refunded automatically. See [Billing](../cloud-api/overview#billing). ## One request end to end ```bash curl -sS https://api.vieneu.io/api/v1/audio/speech \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Content-Type: application/json" \ -d '{"input": "Xin chào, đây là VieNeu.", "response_format": "mp3"}' \ --output speech.mp3 ``` Name `response_format` even though `mp3` is the default. Only an *explicit* mp3 is guaranteed to be mp3: when the request never mentions a format and the worker cannot encode one, the endpoint falls back to wav rather than failing, and `-sS --output` discards the `X-Output-Format` header that would have told you. Written to `speech.mp3`, those are bytes your player refuses. Omitting `voice` uses the first active voice on the configured default engine (see [Engines](../cloud-api/overview#engines) — the engine also sets the price multiplier). To choose one, ask which engine is the default, then list that engine's voices: ```bash curl -sS https://api.vieneu.io/api/v1/engines # no key needed; find "isDefault": true curl -sS "https://api.vieneu.io/api/v1/audio/voices?engine=" \ -H "Authorization: Bearer $VIENEU_API_KEY" ``` `v4` is today's default — and, since `v3` was retired from the cloud API on 2026-09-24, the only engine — but the registry is the source of truth, so substitute whatever the first call reports. The retired catalogue's ids (opaque `vieneu-…` slugs) shared almost no space with `v4`'s display names, so a voice saved from an unfiltered list before the retirement 400s now; pick a fresh one. Details in [Getting voice ids](./openai-clients.md#getting-voice-ids). ## What the OpenAI shape cannot reach `POST /api/v1/audio/speech`, plus the `GET /api/v1/audio/voices` companion its voice picker reads, is the whole compatibility surface — there is no adapter to install and nothing to maintain per client. Two things live outside that shape, and if you are comparing providers they are worth testing directly rather than inferring: - **Incremental streaming.** `POST /v1/tts/stream`, and `stream_format` on the OpenAI route, deliver audio as it is generated rather than as one buffered download wearing a streaming label. It is declared per engine — `features` on `GET /v1/engines` includes `stream` for `v4`; the `v3` engine retired on 2026-09-24 never had it, because it handed over its chunks in a burst at the end and could not honour the latency promise. See [Streaming](../cloud-api/streaming). - **Cloned voices, usable from code.** You clone a voice once in the web Studio ([vieneu.io/#/clone](https://vieneu.io/#/clone)) — the guided path that denoises the clip, transcribes it and plays it back before saving — and its `clone_…` id then works on `POST /v1/tts` and `POST /v1/tts/stream` like any catalogue voice. It is the one thing an OpenAI-compatible client cannot reach: a `clone_…` id on `/v1/audio/speech` is a 400. The smaller frictions of the OpenAI shape — no `GET /v1/models`, unknown body fields rejected rather than ignored, OpenAI voice names unmapped, the doubled `/api/v1` base-URL trap, streaming needing two fields, browser-side CORS, and which 429 came from where — all have a symptom and a fix in [OpenAI-compatible apps](./openai-clients.md). ## Latency, honestly **On the standard plans**, first audio lands in roughly **1–2 seconds**. That is comfortable for read-aloud, IVR prompts, dubbing and assistants that tolerate a beat before speaking. It is **not** the 200–300 ms class that hard real-time conversational agents expect. Enterprise runs on a dedicated node where that ceiling is agreed per contract instead — down to ≤ 250 ms. Either way, design around the number you measure on the plan you are on — see [Latency, honestly](../cloud-api/streaming#latency-honestly). --- # Drop-in for OpenAI-compatible apps Source: https://docs.vieneu.io/docs/integrations/openai-clients A lot of software already speaks `POST /v1/audio/speech` and lets you point it at a custom base URL. VieNeu implements that endpoint, so those apps can use Vietnamese voices with no plugin and no adapter. This page is the setup sheet: what to put in which field, for each client. The VieNeu side of every section below is exact. The client side is written as "which value goes in which kind of field", because setting labels move between versions — **verify the field names against your build**. ## The three values Every client needs the same three things. | Value | What to enter | |---|---| | Base URL | `https://api.vieneu.io/api/v1` (see the trap below) | | API key | your VieNeu key — starts `vn_sk_` (live) or `vn_test_` (test) | | Model | `tts-1`, `tts-1-hd`, `gpt-4o-mini-tts`, `vieneu` or `vieneu-v4` → `v4`, the only engine since `v3` was retired on 2026-09-24. Any other name is accepted and ignored — **except** a `vieneu-…` name that is not a live engine, which is a 400. `vieneu-v3` is now one of those. | The one carve-out in that last row is the one a VieNeu user is most likely to trip. `vieneu-v5` or `vieneu-turbo` does **not** fall through to the default engine the way `whatever-tts` does; it returns 400 `invalid_request_error` with `param: "model"`, listing the engine names that do exist. The prefix is read as an explicit engine choice, and quietly rendering it on some other engine would bill at a rate you never picked. If you send both `engine` and `model`, `engine` wins. `vieneu-v3` is the case you are most likely to actually hit. It was a valid choice until `v3` was retired from the cloud API on 2026-09-24, and now answers 400 with `Model 'vieneu-v3' is retired. Engine "v3" was retired on 2026-09-24. Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues share no ids.` Change the model to `vieneu-v4` (or `tts-1`) and re-check the voice in the same pass — a `vieneu-…` slug saved from the old `v3` catalogue will 400 on `param: "voice"` next. A fourth field, `voice`, is optional but usually wanted: omitting it uses the first active voice on the resolved engine. What you cannot do is reuse OpenAI's names — `alloy` / `nova` / etc. are **not** mapped and return 400. Get real ids from `GET /api/v1/audio/voices` (below). ### The base-URL trap The path is `/api/v1/audio/speech`. The doubled-looking `/api/v1` is real — the server sets a global `api` prefix *and* mounts the public API at `v1`. So the value you type depends on what your client appends to it: | If the field means… | Enter | |---|---| | "the OpenAI base — I append `/audio/speech`" | `https://api.vieneu.io/api/v1` | | "the host — I append `/v1/audio/speech`" | `https://api.vieneu.io/api` | | "the full endpoint URL" | `https://api.vieneu.io/api/v1/audio/speech` | Most clients mean the first. A wrong pick is always a **404, and always JSON** — but in VieNeu's platform error shape rather than the OpenAI envelope this endpoint otherwise uses: ```json { "statusCode": 404, "message": "Cannot POST /api/v1/v1/audio/speech", "error": "Not Found" } ``` Two things to read off it: - The `Cannot POST …` message and the **absence** of the `{"error":{"message","type","param","code"}}` wrapper that every real `/v1/audio/speech` failure carries mean the URL shape is wrong, not the key. Do not go hunting your API key on this one. - The path inside the message is the URL your client actually built. Compare it to `/api/v1/audio/speech` and the difference tells you which row above you needed — the example here doubled `/v1`, so that client wanted `https://api.vieneu.io/api`. ### There is no `/v1/models` VieNeu serves TTS only. There is no `/v1/models` route and no `/v1/chat/completions`. Two consequences: - A **"Test connection" / "Verify" button that probes `/models` will report failure** even though synthesis works. Ignore it and send a real request. - A **model dropdown fed from `/models` will be empty.** Type the model name by hand. In Open WebUI, SillyTavern and LobeChat, this base URL belongs **only** in the audio/TTS provider setting — never in the general OpenAI/LLM endpoint setting, or the chat model breaks. ### Prove it works first Before touching any client, confirm the base URL and the key with curl. Send `$VIENEU_API_KEY` — your own `vn_sk_…` or `vn_test_…` key — and **no `voice`**, so nothing but those two values can fail: ```bash curl https://api.vieneu.io/api/v1/audio/speech \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu." }' \ --output speech.mp3 ``` Omitting `voice` makes the server pick the first active voice on the resolved engine, which is what keeps this step honest: a voice id you have not verified yet would 400 on `param: "voice"` and leave you unable to tell a bad voice from a bad key or a bad base URL. Pick a real voice in the next step, once this one has passed. If that produces playable audio, everything after this is client configuration. ### Getting voice ids ```bash # Which engine will my requests land on? No API key needed. curl https://api.vieneu.io/api/v1/engines # Voice ids for that engine. API key IS required here. curl "https://api.vieneu.io/api/v1/audio/voices?engine=v4" \ -H "Authorization: Bearer $VIENEU_API_KEY" ``` `GET /api/v1/engines` returns one entry per enabled engine with `key`, `isDefault`, `billingMultiplier` and `features`. Use the one with `isDefault: true` — since `v3` was retired on 2026-09-24 that is the only entry, `v4`. :::warning Voice ids saved before 2026-09-24 may be dead `audio/voices` now lists `v4` only, filtered or not, and `?engine=v3` is a 400. But the retired V3 catalogue's ids were `vieneu-…` slugs, V4 ids are display names (`Ngọc Lan`), and the two shared almost no id space — so a voice a client saved from an unfiltered list before the retirement will 400 on `param: "voice"` now. Refresh the picker from the call above, and keep passing `?engine=v4`: it costs nothing and keeps the request explicit. ::: Voice matching is case-insensitive but **not** diacritic-insensitive, and V4 ids contain spaces. A client that slugifies or strips accents before sending will 400 on every request. Now re-run the smoke test with an id from that list, copied exactly as returned: ```bash curl https://api.vieneu.io/api/v1/audio/speech \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu.", "voice": "PASTE_AN_ID_HERE" }' \ --output speech.mp3 ``` If the first curl worked and this one 400s on `param: "voice"`, the id is the only thing that changed — check the engine filter and the accents before anything else. :::note Cloned voices do not work here `clone_…` ids are rejected on `/v1/audio/speech` with a message naming the two routes that accept them: `POST /v1/tts` and `POST /v1/tts/stream`. No OpenAI-compatible client can reach a cloned voice. ::: ## Open WebUI Audio/TTS settings, OpenAI engine: | Field | Value | |---|---| | TTS engine / provider | the OpenAI-compatible option | | API base URL | `https://api.vieneu.io/api/v1` | | API key | `vn_sk_…` | | Model | `tts-1` | | Voice | a real id from `audio/voices`, e.g. `Ngọc Lan` | Notes: - Put this in the **audio** settings, not the model/connection settings. - The voice field must accept free text. If your build only offers OpenAI's six names, they will all 400 — verify on your build. - Response format: leave it at the default. Open WebUI plays mp3, which is also VieNeu's default. ## SillyTavern TTS extension, OpenAI-compatible provider: | Field | Value | |---|---| | Provider | the OpenAI TTS option | | API / base URL | `https://api.vieneu.io/api/v1` | | API key | `vn_sk_…` | | Model | `tts-1` | | Voice (per character) | a real id from `audio/voices` | Notes: - SillyTavern assigns a voice per character. Every one of them must be a VieNeu id — a character left on `alloy` fails while the others work, which reads as a flaky integration. - If the extension only offers a fixed voice dropdown rather than a text field, it cannot address VieNeu voices. Verify on your build. - If your SillyTavern runs its TTS call from the **browser** rather than its own server, see [Browser-side clients](#browser-side-clients). ## LobeChat TTS / audio settings, OpenAI provider: | Field | Value | |---|---| | OpenAI TTS base URL / proxy URL | `https://api.vieneu.io/api/v1` | | API key | `vn_sk_…` | | Model | `tts-1` | | Voice | a real id from `audio/voices` | Notes: - LobeChat keeps separate settings for the chat provider and the TTS provider. This URL goes in the **TTS** one only. - LobeChat is commonly deployed so that the browser calls the TTS provider directly — see [Browser-side clients](#browser-side-clients) before you debug anything else. ## LiteLLM Add VieNeu as a model in `config.yaml`: ```yaml model_list: - model_name: vieneu-tts litellm_params: model: openai/tts-1 api_base: https://api.vieneu.io/api/v1 api_key: os.environ/VIENEU_API_KEY ``` Then, with the proxy running, call it exactly as you would OpenAI: ```bash curl http://localhost:4000/v1/audio/speech \ -H "Authorization: Bearer $LITELLM_MASTER_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "vieneu-tts", "input": "Xin chào", "voice": "Ngọc Lan" }' \ --output speech.mp3 ``` Notes: - **`LITELLM_MASTER_KEY` is not your VieNeu key.** It is LiteLLM's own proxy credential, whatever you set when you started the proxy — the port (`:4000` above) is LiteLLM's too. Your `vn_sk_…` key appears only in `config.yaml`, reached through `os.environ/VIENEU_API_KEY`; the proxy is what attaches it to the upstream call. - The `openai/` prefix on `model` tells LiteLLM to pass the request through in OpenAI's shape, which is what VieNeu answers. The name after it (`tts-1`) is what reaches VieNeu. - The VieNeu values above are exact; the surrounding LiteLLM keys are its standard `model_list` shape — confirm them against your LiteLLM version. - **Do not enable any request-enrichment that adds body fields.** VieNeu rejects unknown fields with 400 rather than ignoring them — see [400](#400-invalid_request_error). ## LiveKit Agents The OpenAI plugin takes a base URL and a key: ```python from livekit.plugins import openai tts = openai.TTS( model="vieneu-v4", voice="Ngọc Lan", base_url="https://api.vieneu.io/api/v1", api_key="vn_sk_...", ) ``` Notes: - **Pin `vieneu-v4`, or omit `model`.** Either lands on `v4`, the only engine since `v3` was retired on 2026-09-24. `vieneu-v3` — which this endpoint used to accept with streaming, answering 200 and delivering v3's chunks in a burst at the end with no error telling you so — is now a 400 with `param: "model"` saying the engine was retired. - `GET /v1/engines` returns the live billing multiplier; see [Engines](../cloud-api/overview#engines). - For streaming, VieNeu needs **both** `response_format: "pcm"` and `stream_format`. `response_format` defaults to mp3, so setting only `stream_format` is a 400. Whether the plugin sends either, and whether it lets you add them, is the thing to verify on your build. Without them the call still works — it just returns a complete mp3 instead of a stream. - `pcm` is headerless and defaults to **24000 Hz** on this route (the engines' native rate is 48000). Read the actual rate from the `X-Sample-Rate` response header and configure the pipeline to match, or the audio plays at the wrong pitch. ## Pipecat ```python import os from pipecat.services.openai.tts import OpenAITTSService tts = OpenAITTSService( api_key=os.environ["VIENEU_API_KEY"], base_url="https://api.vieneu.io/api/v1", model="vieneu-v4", voice="Ngọc Lan", ) ``` Notes: - The import path for Pipecat's OpenAI TTS service has moved between releases — check yours. The constructor arguments and the values above are what matter. - The same three streaming rules as LiveKit apply: pin `vieneu-v4`, send `response_format: "pcm"` **and** `stream_format`, and take the rate from `X-Sample-Rate` rather than assuming 48 kHz. - If your version hard-codes a `response_format` VieNeu does not accept (`aac`, `flac`), the request 400s naming the field. `wav`, `mp3`, `opus`, `pcm` and `ulaw` are the accepted values. ## Browser-side clients In production, VieNeu's CORS policy allows only VieNeu's own origins. A client that calls `/v1/audio/speech` **from the page** rather than from its own server fails at the preflight, with no response body to explain it. Whether a given LobeChat or SillyTavern deployment does its TTS server-side is a per-deployment question — verify on your build. If it is browser-side, put a proxy of your own in front (LiteLLM works well for this) and point the client at that. Browser JavaScript can read `X-Request-Id`, `X-Sample-Rate`, `X-Output-Format`, the `X-RateLimit-*` headers and `Retry-After`; everything else is hidden by the browser. (`X-Stream-Format` is CORS-exposed too, but do not write a read for it here — only the native `POST /v1/tts/stream` sends it. On this route the framing is already unwrapped for you, so the header is always absent.) ## Troubleshooting by symptom ### No audio Work down this list — each has a different cause. - **The saved file contains JSON.** On the raw streaming path, a generation that produced nothing answers **502** with an error body where audio was expected. A client that writes the response body straight to a `.pcm` file ends up with JSON in it. Check the first bytes of the file. - **The file downloads instead of playing.** The synchronous response carries `Content-Disposition: attachment`. Clients that fetch the body are unaffected; one that navigates to the URL gets a download. - **Audio plays at the wrong pitch or speed.** `pcm` and `ulaw` are headerless — the bytes carry no sample rate. `pcm` is 24000 Hz here unless you asked otherwise; `ulaw` is always 8000. Read `X-Sample-Rate` instead of assuming. - **The player refuses an `.mp3`.** If you omitted `response_format` and the worker serving you cannot encode mp3, VieNeu falls back to **WAV bytes** rather than failing. Read `X-Output-Format` to see what you actually got. (Had you asked for mp3 explicitly, you would have got a 503 naming the format instead — an explicit choice is never silently substituted.) ### 400 `invalid_request_error` The `param` field in the error body names the culprit for most of these — with one exception, and it is the first cause listed below. The common causes: - **An unknown body field.** VieNeu rejects fields it does not declare rather than ignoring them. The accepted set is exactly: `model`, `input`, `voice`, `response_format`, `sample_rate`, `speed`, `instructions`, `stream_format`, `emotion`, `aiRefine`, `engine`. Note there is **no `stream` boolean** — a client that sends `stream: true` (as for chat completions) gets a 400. So does `user`, `language`, or any client-specific extension. This is the most common reason a client that works against OpenAI fails against VieNeu: inspect the request body it actually emits. **Read `message`, not `param`, for this one.** The rejection comes from the request validator rather than from the handler, so `param` comes back as the literal string `"property"` and never the offending field name. The name is in the message: `property stream should not exist`. Only the first offending field is reported per request, so if your client adds several, fix and retry until it passes. - **An unrecognised `vieneu-…` model name.** `param: "model"`. Other unknown model names are ignored; this prefix is not. - **An OpenAI voice name.** `alloy`, `echo`, `fable`, `onyx`, `nova`, `shimmer` are not mapped. `param: "voice"`. - **A cloned voice id.** `clone_…` works only on `POST /v1/tts` and `POST /v1/tts/stream`. - **`stream_format` with a non-streamable format.** Only `pcm` and `ulaw` can be streamed, and `response_format` defaults to **mp3** — so setting `stream_format` alone is always a 400. - **`aac` or `flac`.** Real OpenAI formats VieNeu cannot encode; the message says so explicitly. - **A `sample_rate` that contradicts the format.** `opus` is always 48000 and `ulaw` always 8000; a conflicting rate is rejected, not overridden. Valid rates are 8000, 16000, 22050, 24000, 44100, 48000. - **`speed` outside 0.25–4.0.** Inside that range it is clamped to 0.5–2.0 and never errors; outside it, validation rejects it. - **A test key over 100 words.** `vn_test_` keys cap each request at 100 whitespace-separated words. The message names your actual word count. ### 401 `authentication_error` - **`Invalid API key format.`** — the key must literally begin `vn_sk_` or `vn_test_`. This is checked before any lookup, so placeholder keys some clients ship or require (`sk-...`, `none`, `ollama`) fail here. If the client refuses to save an empty key field, it needs a real VieNeu key. - **`API key required.`** although you set one — a bearer token containing a dot is treated as a JWT and ignored as an API key. Real VieNeu keys are hex and never contain a dot, so this means something wrapped or replaced your key. - **`Invalid or revoked API key.`** — a revoked key is indistinguishable from an unknown one. Keys go in `Authorization: Bearer ` or `X-API-Key: `; both work. If your client sends both with different values, `X-API-Key` silently wins. ### 403 — usually "out of credit" {#403-out-of-credit} **Running out of tokens is 403 here, not 402.** Nothing in the public API emits 402, so a client that watches for it never learns it has stopped being paid for. Watch 403 instead. Three messages, all typed `insufficient_quota`: - `Insufficient tokens. You have tokens remaining but this request costs tokens.` — the ordinary end of a grant, and the normal first failure once a trial allowance empties. - `Your API token package has expired. Please renew your Developer plan.` — time, not consumption. Buying more tokens does not fix it; renewing does. - `API token grant is not active. Check your Developer plan status.` Two other 403s are *not* about credit, and the `type` field tells them apart: - `Your plan does not include engine "v4". Upgrade your Developer plan to use it.` — typed `invalid_request_error`, `param: "engine"`. Raised only when you pinned an engine explicitly and your plan does not cover it; with `v4` the only engine since 2026-09-24 you should rarely see it — drop `model`/`engine` and the default engine works. - `This feature is not available yet.` — typed `authentication_error`, from a feature gate rather than from billing. So do not branch on `type` alone to detect "out of credit", and do not branch on 403 alone either. Both together are unambiguous. A daily or weekly cap is **429**, not 403 — see below. ### 429 `rate_limit_exceeded` Read the headers to tell the sources apart: | Headers present | Source | Counted against | |---|---|---| | `Retry-After` + `X-RateLimit-*` | VieNeu's throttler on synthesis routes — 300 requests/min **by default** | your API key | | none of them, non-JSON body | the edge proxy — 3000 requests/min | your **source IP**, pooled with everyone behind it | | JSON, message contains `Resets at ` | your token grant's daily or weekly cap | your grant | Back off in all three cases. The middle one is worth knowing about: if your client runs on shared or NAT'd egress, you can trip a limit no amount of tuning on your side will fix. **Do not hard-code 300.** It is a deployment setting (`PUBLIC_API_SYNTHESIS_RPM`), not a constant, and other pages quote other figures because they were written against other deployments. The row's first column is the durable part: read your live budget from `X-RateLimit-Limit` and `X-RateLimit-Remaining` on any successful response. The third row is the one people mistake for the first. A daily or weekly grant cap is a 429 whose message ends `Resets at ` — no amount of slowing down clears it before that time. Running out of tokens outright is [403](#403-out-of-credit), never 402. ### 503 — the fleet, not your request Nothing about your request needs changing; retry shortly. Four causes: - **No worker for the engine you pinned.** `No V4 TTS worker is available right now. Please try again shortly.` (the streaming path says `…available for streaming right now`.) Raised only for a **non-default** engine: a request that takes the default engine falls back to the configured worker rather than 503, so dropping `model`/`engine` is a valid workaround here. - **No active voice on that engine**, when you omitted `voice`. `No voice is currently available on engine "v4". Pass an explicit voiceId from GET /v1/voices.` This one arrives as a 503 but typed `invalid_request_error` with `param: "voice"` — a catalogue problem wearing a request-shaped label. It is raised before billing, so nothing was charged. - **A format you asked for explicitly that the worker cannot encode.** Had you left `response_format` off you would have got WAV bytes instead; an explicit choice is never silently substituted. See [No audio](#no-audio). - **A `sample_rate` the worker cannot produce**, on the streaming path. The message names both what you asked for and what the worker offered. Every one of these that got as far as billing refunds automatically — you are not charged for audio you did not receive. ### The stream stops mid-sentence - **With SSE, wait for the terminal event.** There is no `data: [DONE]` sentinel. Exactly one `speech.audio.done` (audio complete, carries `usage`) or `speech.audio.error` (it is not) always arrives. If neither did, the stream was truncated — discard the audio; the request is refunded automatically. - **On the raw stream path there is no such signal.** A truncated stream is indistinguishable from a short one. If you need proof of completeness, use `stream_format: "sse"` or the native [`POST /v1/tts/stream`](../cloud-api/streaming). - **It "hangs", then dumps everything at once.** That used to be `vieneu-v3` with `stream_format`: the endpoint accepted it, returned 200, and v3 emitted its chunks in a burst at the end. Since `v3` was retired on 2026-09-24 that request is a 400 instead, so if you still see this pattern the buffering is happening in a proxy or client of your own, not on the engine. Pin `vieneu-v4` and check what sits between you and the API. - **A non-streaming call times out on long text.** `input` has **no length cap** on this endpoint — the request simply blocks for the full synthesis time, and the client's own HTTP timeout gives up first. The ceilings are a 2 MB request body and a 600s proxy read timeout. For long text use the asynchronous `POST /api/v1/tts` + `GET /api/v1/tts/{jobId}` path instead. Every response carries `X-Request-Id`. Quote it in a support request. ## See also - [Integrations overview](./overview.md) — which route to take if your tool is not on this page. - [OpenAI-compatible TTS endpoint](../cloud-api/openai-compatible) — the full field-by-field contract for `/v1/audio/speech`: every accepted field, the format and sample-rate rules, and the complete error list. - [Streaming](../cloud-api/streaming) — both streaming endpoints compared, and the terminal-event semantics summarised above. - [Rate limits and request ids](../cloud-api/overview#rate-limits-and-request-ids) — the headers, and why limits count per key. - [API reference](/api-reference) — generated from the server's own OpenAPI spec. --- # n8n node Source: https://docs.vieneu.io/docs/integrations/n8n :::info Release status: built, not yet published The VieNeu community node is finished and tested, but **the package has not been released publicly**. It is marked `private: true` in its `package.json` and has never been pushed to npm, so there is nothing for you to install today from a public registry. If you want to run it now, **contact VieNeu** and ask for the package — we can hand you a build. Watch the [Cloud API changelog](../cloud-api/changelog) for the announcement when it reaches npm; the [Install](#install) section below has both paths. Everything else on this page — the operations, parameters, output fields and error behaviour — is accurate against the built package, so you can decide whether the node is what you want before asking for it. ::: VieNeu ships an n8n community node — one node, four resources, audio out as a binary field. It is a **programmatic node**, not a trigger: it sits in the middle of a workflow and turns Vietnamese text into a file the next node can send, store or upload. There are two ways to call VieNeu from n8n: - **The node** — a voice picker with search, automatic routing between the short and long text endpoints, and API errors mapped onto n8n's machine-readable `failure.cause`. Needs the package (see above) and a self-hosted n8n. - **Plain HTTP Request nodes** — nothing to install, works on n8n Cloud, and available to everyone today. See [Don't want to install the node?](#dont-want-to-install-the-node) at the bottom. Two importable templates ship with the package. :::caution Do not `npm install n8n-nodes-vieneu` The name is **unclaimed on npm**, so whatever that command resolves to today is not this package. Do not install it and do not paste it into n8n's community-node dialog. ::: ## Install Self-hosted n8n only. n8n Cloud installs verified packages from npm, so the community-node path is unavailable there whatever happens — use the [HTTP Request route](#dont-want-to-install-the-node) instead. ### When the package is published This is the shape the install will take. **None of it works yet** — the package is not on npm: ```text Settings → Community nodes → Install → n8n-nodes-vieneu ``` Nothing below will be needed then. Check the [changelog](../cloud-api/changelog) before following the manual path. ### Today Ask VieNeu for the package. What arrives is the source directory `n8n-nodes-vieneu/`, which you build yourself; the repository it lives in is private, so there is no `git clone` to give you. Prerequisites: | Requirement | Where it comes from | | --- | --- | | **Node 20.19 or newer** | `engines.node: ">=20.19"` in the package's `package.json` | | **pnpm** — `corepack enable` is enough | the package pins `packageManager: "pnpm@10.24.0"`; npm and yarn are not substitutes here (see below) | | **n8n bundling `n8n-workflow` 2.37.1 or newer** | the version the package is compiled against, and the source imports `NodeConnectionTypes`, which older releases do not export | n8n does not publish a plain "minimum version" number for this; check what your instance actually bundles with `npm ls n8n-workflow` where n8n is installed. The package declares `n8nNodesApiVersion: 1`. On an n8n too old to export `NodeConnectionTypes` the node does not appear in the panel at all; on one too old for the `Failure` type, the `failure.cause` branching documented [below](#how-failures-surface) does nothing. Build it: ```bash cd n8n-nodes-vieneu pnpm install pnpm build pnpm verify ``` `pnpm build` is mandatory. `dist/` is git-ignored, and `dist/` is exactly what the `n8n` block in `package.json` points n8n at — a fresh copy has none. `pnpm verify` is the check worth running. It `require()`s the built `dist/` the way n8n's loader does and instantiates the exported class, which catches the three failures that otherwise show up only as a node that silently never appears in the panel: a `package.json` path that no longer matches the emitted layout, an icon `tsc` did not copy, and a class n8n cannot construct. Use `pnpm`, not `npm install`. The package sits outside the repository's own workspace and carries an empty `pnpm-workspace.yaml` to make its directory a workspace root; without pnpm reading that file, an install run from here walks up to the repository root and installs nothing at all, exiting 0 with nothing to build from. Then either copy the build in, or point n8n at the package directory: ```bash # Option A — copy into n8n's private-node folder, then restart n8n. mkdir -p ~/.n8n/custom cp -r dist/nodes dist/credentials ~/.n8n/custom/ # Option B — leave it in place and point n8n at it. export N8N_CUSTOM_EXTENSIONS="/abs/path/to/n8n-nodes-vieneu" n8n start ``` Two mistakes that cost an afternoon: - `~/.n8n/custom` is **not** `~/.n8n/nodes`. The latter holds npm-installed community nodes; a private build placed there is ignored. - `N8N_CUSTOM_EXTENSIONS` takes a **semicolon**-separated list of absolute paths on every platform. A colon-separated list does not error — it loads nothing. For Docker, mount the built package into the container and set that same variable to the in-container path. The node should then appear as **VieNeu** in the node panel. Its full type is `n8n-nodes-vieneu.vieneu`. :::caution This install path has not been exercised against a live n8n The package compiles, its unit tests pass, and `pnpm verify` loads the built output in the shape n8n's loader expects — but nobody has yet loaded it into a running n8n instance, and the same is true of the two workflow templates, which are validated structurally rather than by importing them. Treat your first install as a smoke test, and tell us what breaks rather than assuming it is something you did. ::: :::note Not an AI Agent tool The node deliberately does not declare `usableAsTool`, so an AI Agent cannot call it. Speech generation spends the account's tokens on every call, and an agent invoking it speculatively is the wrong first experience. Agents have their own supported surface: the [MCP server](./mcp/index.md). ::: ## Credential One credential type, **VieNeu API**, with two fields and nothing else. | Field | Name in workflow JSON | Required | Default | | --- | --- | --- | --- | | API Key | `apiKey` | Yes | — (password-masked, placeholder `vn_sk_…`) | | Base URL | `baseUrl` | Yes | `https://api.vieneu.io/api/v1` | **The `/api/v1` prefix in the Base URL is real.** The controller is mounted at `v1` but the application sets a global `api` prefix, so the documented `/v1/...` paths are served at `/api/v1/...`. A base URL of `https://api.vieneu.io/v1` 404s on everything. Change this field only to point at staging or a self-hosted deployment; trailing slashes are trimmed, and an empty value falls back to the default. Press **Test** after pasting a key. The test issues `GET /voices` against the configured base URL — it costs no tokens and is exempt from the rate limiter, but its guard still rejects a malformed or revoked key with 401, so a wrong key fails here instead of quietly returning the public catalogue. Which keys are valid, and what a `vn_test_` key can do, is the same everywhere: see [Authentication](../cloud-api/overview#authentication). The one thing that matters in n8n is that a `vn_test_` key's 100-word-per-request cap surfaces as a **400** on long text, not as a quota error — the node does not check the prefix or the word count locally. ### The key never reaches a workflow variable This is the reason the key is a credential rather than a node parameter: - Node parameters are written into execution history, into the workflow JSON people paste into issues, and into log lines. A credential is encrypted at rest and referenced by id — it is not part of an exported workflow. - No code in the package ever reads `apiKey`. n8n builds `Authorization: Bearer {{$credentials.apiKey}}` itself, inside `httpRequestWithAuthentication`, from the credential's `authenticate` block. There is no variable holding the key that an expression or a log line could reach. - Every error message, description and attached response body is run through a redactor first. Bearer tokens and anything shaped like `vn_sk_…`, `vn_test_…` or `sk_…` become `***REDACTED***` before n8n stores them in execution data. ## Resources and operations | Resource | Operation | Calls | Parameters | | --- | --- | --- | --- | | Speech | Generate | `POST /audio/speech` or `POST /tts` + poll | Text, Engine, Voice, Put Output File in Field, Options | | Job | Get | `GET /tts/{jobId}` | Job ID, Download Audio, Put Output File in Field | | Voice | Get Many | `GET /voices` | Engine, Search, Return All, Limit | | Engine | Get Many | `GET /engines` | none | Engine: Get Many has no parameters of its own. It returns one item per engine, including the **live** `billingMultiplier` — read that rather than any table, including the snapshot in [Engines](../cloud-api/overview#engines). ### What the node does not do Not everything the API offers is exposed, and the omissions are deliberate: an operation that spends real money or needs a file upload is a bad thing for a workflow to reach by accident, and a Loop Over Items reaches things many times. | Missing | Why | | --- | --- | | **Voice cloning** (`POST /voices`, `/clone`, `/prepare`, `/upload`) | A flat 5,000-token charge before any multiplier (15,000 on v4) for one clip, a multipart upload, and a slot against the plan's clone limit. Clones you already own *are* selectable in the Voice list; you just cannot create one from a workflow. | | **`DELETE /voices/:voiceId`** | Destructive, and irreversible from inside a workflow. | | **Dubbing and SRT** (`/dub`, `/srt`) | Both upload-shaped, both billed per call. | | **`POST /dialogue`** | Key-only and no upload, but 50 turns × 5,000 characters is uncapped in the request schema — by a wide margin the largest single-call exposure on the surface. | | **Streaming** (`POST /tts/stream`) | Its response is a custom length-prefixed frame format with heartbeat frames mixed in; n8n has no frame parser and would hand back one opaque buffer. It also caps at 4 concurrent streams per key, and every stream bills at the v4 multiplier whatever engine was sent. See [Streaming](../cloud-api/streaming) if you need that surface. | ### Picking a voice The Voice field is a resource locator with two modes: - **From List** — searchable, backed by `GET /voices`. Filtering happens in the node, because the API has no search parameter: it folds diacritics (`đ`/`Đ` included) and requires every word to match against id, name, description, gender, region, engine and kind. Rows read `Name — gender · region · engine (id)`, with `· your clone` inside the facet group for **any** voice the API returns as `kind: "cloned"` — your own clones and admin-published ones alike, so a public clone somebody else created is labelled that way too. The list pages 250 at a time. - **By ID** — for expressions. Copy the id exactly, diacritics included. Setting **Engine** first scopes the list, and on Speech: Generate the list re-loads when Engine changes. Do that: a voice renders only on its own engine and is never substituted — a mismatch is a hard 400. The dropdown is loaded live from `GET /engines`, so since `v3` was retired on 2026-09-24 it offers `v4` alone; the only mismatch left is an old `v3` slug pasted By ID. :::warning A `clone_…` voice does not work on short text The short-text route the node picks by default, `POST /audio/speech`, rejects every cloned voice outright — its voice check runs without a user id, so it cannot tell your clone from anyone's and refuses them all with a 400 reading *"Cloned voice … can only be used on POST /v1/tts or POST /v1/tts/stream."* This applies to admin-published clones too. The long-text route is `POST /tts`, which does accept clones — your own and published ones. So to use a `clone_…` voice from the node, keep the item **over** the Synchronous Route Threshold, or set **Options → Synchronous Route Threshold** to `1` to force every item onto the async route. Do not follow the node's 400 advice here. It says the usual cause is a voice belonging to the other engine, which is true of catalogue voices and has nothing to do with this. ::: Ids are not descriptive. Many are opaque slugs such as `vieneu-2-000494` whose voice is named `Minh Đức`, and a few are an unrelated person's name outright. Never infer gender or identity from an id. ## Speech: Generate | Parameter | Type | Default | Notes | | --- | --- | --- | --- | | Text | string (4 rows) | — | Required. Billed per character with a 50-character minimum — see [Billing](../cloud-api/overview#billing). | | Engine Name or ID | options | account default | Empty = the account default decides, which also decides the rate. | | Voice | resource locator | engine default | Scoped by Engine. | | Put Output File in Field | string | `data` | Required. The binary property the audio lands in. | | Options | collection | `{}` | Nine options, below. | ### Options | Option | Name | Default | Notes | | --- | --- | --- | --- | | AI Refine | `aiRefine` | `false` | Runs AI normalisation and moderation first. Costs a surcharge on top of the character count and adds latency. No `/v1` route reports the surcharge — the node's help text cites 1.3× at the time of writing. It is also the only thing that can produce a 422. | | Emotion Name or ID | `emotion` | engine default | Loaded from `GET /emotion-tags` for the selected engine. Engine-specific: v4 has no styles, so the list stays empty there. | | File Name | `fileName` | derived | Empty derives a name from the format the API actually produced. | | Format | `format` | `wav` | `wav`, `mp3`, `opus`, `pcm`, `ulaw`. | | Download Audio | `downloadAudio` | `true` | Long-text route only — see below. | | Job Timeout (Seconds) | `jobTimeout` | `600` | Minimum 5. How long to poll a long-text job. | | Poll Interval (Seconds) | `pollInterval` | `2` | Minimum 0.5, and clamped to 500 ms in code regardless of what the workflow JSON says. | | Speed | `speed` | `1` | 0.5–2.0. | | Synchronous Route Threshold | `syncMaxChars` | `500` | Minimum 1. Where the route split happens. | The node sends no sample rate and exposes no option for one. :::tip You may not need Job Timeout at all Job Timeout exists because the node polls. If your jobs are long enough that the timeout is a real risk, register a webhook instead: VieNeu POSTs a signed event when the job reaches a terminal state, and an n8n **Webhook** trigger node receives it. `POST /tts` — the route this option governs — is the only route that produces those events. See [Webhooks](../cloud-api/webhooks). ::: ### Two routes, chosen by text length | | Short text | Long text | | --- | --- | --- | | Condition | length ≤ `syncMaxChars` (default 500) | longer | | Endpoint | `POST /audio/speech` (blocking) | `POST /tts`, then poll `GET /tts/{jobId}` | | Body fields | `input`, `response_format`, `speed`, and `voice` / `engine` / `emotion` / `aiRefine` when set | `text`, `speed`, and `voiceId` / `engine` / `emotion` / `aiRefine` when set | | Formats | all five | WAV only — the route has no format parameter | | Cloned voices | rejected with 400 | accepted (yours and admin-published) | | `Download Audio` | ignored; bytes come back inline | applies | | Extra JSON out | `sampleRate`, `requestId` | `jobId`, `duration`, `audioUrl`, `audioUrlExpiresIn` | `format` and `mimeType` are set on **both** routes and are listed with the shared fields [below](#output) — do not treat their presence as a signal of which route an item took. Read `route` for that. Note the field names differ between the routes (`input`/`voice` versus `text`/`voiceId`). The node handles that; it matters only if you copy a body between the node and an HTTP Request node. Raising `syncMaxChars` makes the node hold an HTTP connection — and an n8n worker slot — open for the whole synthesis. The finished audio on the long-text route is fetched from a presigned S3 URL **without** authentication. S3 rejects a presigned request that also carries an `Authorization` header, so that one call deliberately skips the credential. ### Checks that run before anything is billed These fail locally, with no request sent and no money spent: - empty text; - text over 50,000 characters; - speed outside 0.5–2.0; - a non-WAV format combined with text over the threshold. The long-text route can only return WAV, so this is refused rather than silently downgraded. ### Output Every item's JSON carries: | Field | Meaning | | --- | --- | | `route` | `sync` or `async` — the only reliable way to tell which route ran | | `engine`, `voiceId` | what was requested, or `null` for the default | | `textLength` | raw character count | | `billedCharacterBasis` | NFC character count with a 50-character floor | | `billingNote` | states that the basis is before the engine multiplier and the AI-refine surcharge | | `format` | the format actually produced; always `wav` on the long-text route | | `mimeType` | read off the response, not the request — see below | `billedCharacterBasis` is the **basis** of the charge, not a token total. The backend applies four factors in this order: the character basis (50-character floor), then the AI-refine surcharge if it is on, then the engine multiplier, then a **global token multiplier** — an administrator setting, 1 by default, that no `/v1` route reports. That last factor is why no number the node prints is a quote. Read the engine multiplier live from Engine: Get Many, and treat any budget you compute as approximate. ## Where the audio comes out The audio is attached to the binary property named by **Put Output File in Field** — `data` unless you change it. The same item's JSON also gains `fileName` and `fileSize`. Rename the field when a node upstream already occupies `data`, or when a downstream node expects a specific name. The mime type and file name are read off the **response**, never from the request: 1. the `X-Output-Format` header, if it names a known format; 2. otherwise the `Content-Type`; 3. the file name comes from `Content-Disposition` when present, else `speech.` — `wav`, `mp3`, `ogg` for Opus, and `raw` for the headerless `pcm` and `ulaw`. That indirection is deliberate: the short-text route can fall back to WAV against an older worker, and these headers are the only signal that it did. Labelling the binary from the request would hand a downstream node WAV bytes marked `audio/mpeg`. On the long-text route the default name is `speech-.wav`. **Feeding the next node.** Any node that consumes a file asks for the binary property by name: give it `data`, or whatever you renamed the field to. The size and name are also on the item's JSON as `{{ $json.fileSize }}` and `{{ $json.fileName }}`, so a downstream IF can check them without touching the bytes. To keep a large WAV out of workflow memory when you only need the link, switch **Download Audio** off on the long-text route and pass `{{ $json.audioUrl }}` along instead — it stays valid for a limited time, and the JSON reports how long as `audioUrlExpiresIn`. ## How failures surface Every API failure becomes a `NodeApiError` whose message reads: ```text VieNeu API error 403 while synthesizing speech — (request id ) ``` The prose is in `description`. The part to branch on is n8n's machine-readable `failure.cause`. What each status means on the API side is in [Errors](../cloud-api/overview#errors); this table is the mapping the node adds on top: | Status | `failure.cause` | What happened | | --- | --- | --- | | 400 | `configuration-invalid` | Usually a voice belonging to the other engine — but also a `clone_…` voice on the short-text route, non-Vietnamese text, and a `vn_test_` key over its 100-word cap. | | 401 | `credential-invalid` | The key was rejected. Paste a current one into the credential. | | 403, message mentions an engine | `configuration-invalid` | The plan does not include that engine. | | 403, otherwise | `quota-exhausted` | The token grant is exhausted or expired. **No wait hint is set** — it does not clear on its own. | | 404 | `configuration-invalid` | Nothing under that id. A jobId from another key reads exactly like one that never existed; a withdrawn voice answers 404, not 400. | | 413 | `configuration-invalid` | Payload too large. | | 422 | `configuration-invalid` | Moderation refused the text. Only reachable with AI Refine on. Rephrase; retrying unchanged fails again. | | 429 | see below | Three different limiters. | | 503 | `temporarily-unavailable` | No worker free for that engine. Transient. | | other 5xx | `temporarily-unavailable` | Retry shortly. | | anything else | `node-defect` | Unexpected response. | VieNeu answers **403** when tokens run out — never 402. Retry logic keyed on 402 will never fire. DNS, TLS, reset and timeout failures never reach that table; they become a `temporarily-unavailable` error reading "Could not reach VieNeu while …", whose description tells you to check that the Base URL includes `/api/v1`. ### Rate-limited versus quota-exhausted A 429 can come from three places that want opposite responses. The node reads the response headers to tell them apart — which is why it inspects the status itself rather than letting n8n throw and lose them. | Signal on the response | Which limiter | `failure.cause` | Wait hint | What to do | | --- | --- | --- | --- | --- | | `Retry-After` present | The app throttler, counted **per API key** | `rate-limited` | `retryAfterMs`, from the header | Wait the stated time and retry. Nothing is wrong with the balance. | | No `Retry-After`, body matching `token limit reached` / `resets at` / `quota` | The **token quota** — a daily or weekly cap on the grant | `quota-exhausted` | `resetsAtEpochMs`, parsed out of the message text | Do not retry. Wait for the named reset, or top up. | | No `Retry-After`, no such wording (nginx HTML) | The **nginx edge**, counted per source IP | `rate-limited` | none | Back off hard. On n8n Cloud that IP is shared with unrelated customers. | Two of the three collapse onto `rate-limited`. Tell them apart by whether `retryAfterMs` is set: present means the app throttler and a known wait; absent means the edge and no wait hint at all. The quota reset timestamp is parsed out of prose because it has to be: the controller keeps only the message string, so the text is the last surviving trace of the structured `resetAt` field. Note that two different statuses map to `quota-exhausted`: the 429 above, which carries a reset, and the 403 grant exhaustion, which carries none. ### Failures that are not HTTP failures - **A failed job returns HTTP 200** with `status: "failed"`. Speech: Generate raises it as an error; Job: Get returns it as data, so a Wait + IF loop can see the terminal state and leave the loop. - **A poll timeout** raises a 504-shaped error stating that the job is still running and has **already been billed**, naming its `jobId`. Collect it later with Job: Get. Keep Job Timeout under n8n's own execution timeout, or n8n kills the run and the jobId goes with it. - **An empty audio body**, or a submit that returns no `jobId`, raises a 502-shaped error. - **A 200 with an empty voice list and an `error` field** — the catalogue read failing — is raised as an error rather than rendered as an empty dropdown. When the node is set to continue on failure instead of stopping, the failure becomes item JSON: `error` with the redacted message, plus `jobId` **only when a job was already submitted**. Nothing else is on that item — in particular, no binary. A short-text failure, and an async failure that happened at submit, carry `error` with no `jobId` at all. ## Example workflow: narrate long text, recover the ones that time out Seven minutes of audio can outlast a poll window, and the job is billed at submit. This workflow keeps those jobs instead of paying for nothing. If you can receive an inbound HTTPS request, a [webhook](../cloud-api/webhooks) and an n8n **Webhook** trigger node avoid the whole problem — there is no poll window to outlast. Build this only when you cannot. 1. **Manual Trigger** — "When clicking 'Execute workflow'". Replace it with whatever produces your items; each item needs a `text` field. 2. **VieNeu** — Resource *Speech*, Operation *Generate*. - **Text**: `={{ $json.text }}` - **Engine**: pick one, so the voice list is scoped to it. - **Voice**: *From List*, search for the voice you want. - **Put Output File in Field**: `data` - **Options → Format**: `WAV` (long text returns nothing else). - **Options → Job Timeout (Seconds)**: `900`, below your n8n execution timeout. - Open the node's **Settings** tab and change **On Error** from *Stop Workflow* to the option that continues using the node's **regular output** — not the one that adds a separate error output connector, which would send failed items down a branch the IF in step 3 never sees. Without this, one timed-out job aborts the whole run and the jobId is lost. 3. **IF** — "Failed?". Condition: `={{ $json.error }}` *is not empty*. True means something went wrong; false means the item has its audio on `data`. 4. **True branch → IF** — "Recoverable?". Condition: `={{ $json.jobId }}` *is not empty*. True means the job was billed and is probably still running. False means the failure happened before a job existed — a bad key, a wrong voice, a 429 — so there is nothing to collect: send it to an error output, a Slack message, or wherever you want to see it. Do **not** merge it into step 8; it has no binary and never will. 5. **Recoverable → Wait** — 60 seconds. Polling costs no tokens and no rate-limit budget, but it is not free at the edge. 6. **Wait → VieNeu** — Resource *Job*, Operation *Get*. - **Job ID**: `={{ $json.jobId }}` - **Download Audio**: on (it is off by default on this operation) - **Put Output File in Field**: `data` 7. **IF** — "Terminal?". Condition: `={{ $json.status }}` equals `completed`. True → step 8. False → an IF on `={{ $json.status }}` equals `failed`, whose true side ends the branch (a No Operation node) and whose false side loops **back to the Wait node** in step 5. A job that outlasted a 900-second poll is not reliably finished 60 seconds later, and without this loop the item falls through with no binary attached. 8. Merge the completed items — those from step 3's false branch and those from step 7's true branch — into whatever consumes the file: upload, email, storage. Both carry the same binary property name, so downstream nodes do not need to know which path an item took. The same shape works item-by-item over a spreadsheet. Remember that each row is a separate billed synthesis. ## Don't want to install the node? Every operation is one HTTPS call, and n8n's built-in **HTTP Request** node makes all of them. This path works on n8n Cloud, needs no build step, and needs nothing from VieNeu but a key. Two ready-made workflows ship with the package, under `n8n-nodes-vieneu/templates/`: | File | What it builds | | --- | --- | | `vieneu-speech-http-request.json` | Manual trigger → one HTTP Request to `POST /audio/speech` → MP3 on the `data` binary field. | | `vieneu-long-text-http-request.json` | Submit to `POST /tts`, then Wait → poll → IF until the job is terminal, then download the presigned audio. | Import either with *Workflows → Import from File*. Both attach the key through a Bearer Auth credential, never a header typed into a node parameter. Note that neither has been executed against a live n8n — they are validated structurally. If you do not have the package, the recipes below build the same two workflows by hand. For the underlying request contract, see the [Cloud API overview](../cloud-api/overview). Find a voice id first. Both calls below need no key. Since 2026-09-24 the cloud API has one engine, V4, whose ids are display names (`Ngọc Lan`); the `vieneu-…` slugs of the retired V3 catalogue no longer resolve and 400 on `voiceId`, so refresh any id you saved before then. ```bash # Which engine is the default? Look for "isDefault": true. curl -s https://api.vieneu.io/api/v1/engines # Then list that engine's voices only. curl -s "https://api.vieneu.io/api/v1/voices?engine=v4" | head -c 400 ``` Check that your key works, and hear the result, before wiring anything: ```bash curl -X POST https://api.vieneu.io/api/v1/audio/speech \ -H "Authorization: Bearer $VIENEU_API_KEY" \ -H "Content-Type: application/json" \ -d '{"input":"Xin chào Việt Nam.","response_format":"mp3","speed":1}' \ --output speech.mp3 ``` ### Short text: one HTTP Request node | Setting | Value | | --- | --- | | Method | `POST` | | URL | `https://api.vieneu.io/api/v1/audio/speech` | | Authentication | Generic Credential Type → **Bearer Auth**, holding your `vn_sk_…` key | | Send Body | on, JSON | | Response → Format | **File**, output property `data` | Body, as an expression so the text can come from the incoming item: ```js {{ JSON.stringify({ input: $json.text || 'Xin chào Việt Nam! Đây là giọng đọc tiếng Việt của VieNeu.', voice: $json.voiceId || undefined, response_format: 'mp3', speed: 1 }) }} ``` Use a **Bearer Auth credential**, not an `Authorization` header typed into the node. Credentials are referenced by id and are not part of the exported workflow JSON — that is the difference between a workflow you can paste into an issue and one that leaks your key when you do. `response_format` accepts `mp3`, `wav`, `opus`, `pcm`, `ulaw` here. This route blocks for the whole synthesis, so keep it for short text. A `clone_…` voice is rejected here with a 400, same as through the node. ### Long text: submit, poll, download Past roughly 500 characters, queue the job instead of holding a connection open. Before you build the loop: if your n8n can receive an inbound HTTPS request, a [webhook](../cloud-api/webhooks) replaces steps 2 to 5 with a single **Webhook** trigger node. VieNeu POSTs a signed event the moment the job reaches a terminal state, and `POST /tts` is the only route that produces those events. Poll only when an inbound endpoint is not available to you. The polling version, matching `vieneu-long-text-http-request.json`, is eight nodes: a Manual Trigger and the seven below. 1. **HTTP Request — "Submit TTS job"**: `POST https://api.vieneu.io/api/v1/tts`, Bearer Auth credential, JSON body: ```js {{ JSON.stringify({ text: $json.text, voiceId: $json.voiceId || undefined, speed: 1 }) }} ``` Note the field names on this route: `text` and `voiceId`, not `input` and `voice`. It takes no format parameter and always returns WAV. It responds with a `jobId`, and the text is billed **here**, before the job is queued. 2. **Wait** — 3 seconds. 3. **HTTP Request — "Check job status"**: `GET https://api.vieneu.io/api/v1/tts/{{ $json.jobId }}` (as an expression), same Bearer Auth credential. 4. **IF — "Completed?"**: `={{ $json.status }}` equals `completed`. True → download. False → step 5. 5. **IF — "Failed?"**: `={{ $json.status }}` equals `failed`. True → step 6. False → back to the **Wait** node. A failed job answers HTTP **200**, so branch on the body, never the status code. 6. **No Operation — "Job failed"**: ends the failed branch. 7. **HTTP Request — "Download audio"**: `GET {{ $json.audioUrl }}`, Response Format **File**, output property `data`, and **no credential and no authentication**. The URL is a presigned S3 link carrying its own signature — S3 rejects the request outright if an `Authorization` header rides along. Errors on this path arrive raw, without the node's classification. The rules still hold: a 429 with `Retry-After` is the per-key throttler; a 429 whose message mentions a token limit or a reset time is the quota; a 429 with neither came from the edge and is counted per source IP. Running out of tokens entirely is a 403. --- # Installation Source: https://docs.vieneu.io/docs/getting-started/installation ## Prerequisites - **Python 3.10+** - **eSpeak NG** — Required for phonemization ### Install eSpeak NG ```bash # macOS brew install espeak # Ubuntu/Debian sudo apt install espeak-ng # Fedora/Amazon Linux sudo dnf install espeak # Windows # Download .msi from https://github.com/espeak-ng/espeak-ng/releases ``` ### Optional: NVIDIA GPU For maximum speed via LMDeploy or GGUF GPU acceleration: - NVIDIA Driver >= 570.65 (CUDA 12.8+) - [NVIDIA GPU Computing Toolkit](https://developer.nvidia.com/cuda-downloads) ## Install from Source (Recommended) ```bash git clone https://github.com/pnnbao97/VieNeu-TTS.git cd VieNeu-TTS ``` ### GPU Support (Default) ```bash uv sync ``` ### CPU-Only (Lightweight) ```bash # Linux/macOS cp pyproject.toml pyproject.toml.gpu cp pyproject.toml.cpu pyproject.toml uv sync ``` ## Install as Python Package ```bash # Windows (CPU optimized) pip install vieneu --extra-index-url https://pnnbao97.github.io/llama-cpp-python-v0.3.16/cpu/ # macOS (Metal GPU accelerated) pip install vieneu --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/metal/ # Linux / Generic pip install vieneu ``` ## Verify Installation ```python from vieneu import Vieneu tts = Vieneu() audio = tts.infer(text="Xin chào") tts.save(audio, "test.wav") print("Installation successful!") ``` --- # Quick Start Source: https://docs.vieneu.io/docs/getting-started/quick-start ## Web UI The fastest way to try VieNeu-TTS: ```bash uv run vieneu-web ``` Open `http://127.0.0.1:7860` — type text, pick a voice, click generate. ## Real-time Streaming For ultra-low latency streaming (CPU optimized): ```bash uv run vieneu-stream ``` Open `http://localhost:8001` — audio starts playing before the sentence finishes. ## Python SDK ### Basic Usage ```python from vieneu import Vieneu tts = Vieneu() # Generate speech with default voice audio = tts.infer(text="Xin chào, tôi là VieNeu.") tts.save(audio, "output.wav") ``` ### Voice Cloning ```python audio = tts.infer( text="Đây là giọng nói được clone.", ref_audio="path/to/reference.wav", ref_text="Transcript of the reference audio." ) tts.save(audio, "cloned.wav") ``` ### Using Preset Voices ```python # List available voices voices = tts.list_preset_voices() for description, voice_id in voices: print(f"{voice_id}: {description}") # Use a specific voice voice = tts.get_preset_voice("voice_name") audio = tts.infer(text="Chào bạn!", voice=voice) ``` ### Streaming ```python for audio_chunk in tts.infer_stream(text="Một đoạn văn dài..."): # Process each audio chunk as it's generated play_audio(audio_chunk) ``` ### Batch Processing ```python texts = ["Câu một.", "Câu hai.", "Câu ba."] audios = tts.infer_batch(texts) for i, audio in enumerate(audios): tts.save(audio, f"output_{i}.wav") ``` --- # System Requirements Source: https://docs.vieneu.io/docs/getting-started/system-requirements ## Minimum Requirements | Component | Requirement | |-----------|------------| | Python | 3.10+ | | RAM | 2 GB (GGUF Q4) | | Disk | ~500 MB (model auto-downloaded) | | eSpeak NG | Required | ## Recommended (CPU) | Component | Recommendation | |-----------|---------------| | CPU | Modern i5/i7/M1+ | | RAM | 4 GB+ | | Model | GGUF Q4 or Q8 | Streaming latency: under 300ms first chunk on modern i3/i5. ## Recommended (GPU) | Component | Recommendation | |-----------|---------------| | GPU | NVIDIA with 4GB+ VRAM | | Driver | >= 570.65 (CUDA 12.8+) | | Model | PyTorch 0.5B or 0.3B | | Backend | LMDeploy for maximum speed | ## Intel Arc GPU Supported via PyTorch XPU (tested on Arc B580, A770 on Windows): ```bash run setup_xpu_uv.bat run run_xpu.bat ``` Tip: Intel Arc has high memory bandwidth — keep batch size high and minimize characters per chunk. ## Model Sizes | Model | Disk | RAM Usage | |-------|------|-----------| | 0.5B PyTorch | ~2 GB | ~3 GB | | 0.3B PyTorch | ~1.2 GB | ~2 GB | | 0.3B GGUF Q8 | ~350 MB | ~500 MB | | 0.3B GGUF Q4 | ~200 MB | ~300 MB | Models are cached at `~/.cache/huggingface/hub/` after first download. --- # SDK Overview Source: https://docs.vieneu.io/docs/sdk/overview The `vieneu` package runs **VieNeu-TTS v3 Turbo** on your own machine. It defaults to the torch-free ONNX engine on CPU and switches to PyTorch on a CUDA GPU with no code change. :::tip Where does this fit? The SDK is the **on-device** path — free, open source (Apache 2.0), your hardware. If you would rather call a hosted API (including the proprietary **v4** engine with higher cloning fidelity), see the [Cloud API](/docs/cloud-api/overview). ::: The pages after this one go deeper on one topic each: - [Install & backends](/docs/sdk/standard-mode) — `pip install vieneu`, CPU vs GPU, precision, v3 Nano - [GPU batching](/docs/sdk/fast-mode) — `infer_batch`, CUDA graphs, throughput numbers - [Streaming](/docs/sdk/streaming) — `infer_stream`, concurrent streams - [Voice cloning](/docs/sdk/voice-cloning) — `ref_audio`, `add_voice`, `denoise` - [OpenAI-compatible server](/docs/sdk/remote-mode) — `/v1/audio/speech` from the repo or Docker, plus the legacy v2 `remote` mode :::info Source This section mirrors the **Using the Python SDK** part of the open-source [README](https://github.com/pnnbao97/VieNeu-TTS#readme) and is refreshed automatically (last sync 2026-09-16). If something here disagrees with the README, the README wins — [open an issue](https://github.com/pnnbao97/VieNeu-TTS/issues) there. ::: The `vieneu` SDK **defaults to VieNeu-TTS v3 Turbo (48 kHz)**. The minimal install is **torch-free**: on CPU everything runs on **ONNX Runtime** (PyTorch is never imported), and on a CUDA machine it auto-switches to the PyTorch engine — where inference is **batched automatically** (same API, no code change). ## Quick Start **CPU (default)** — torch-free, runs v3 Turbo via ONNX Runtime. Most users want this: > ⚡**On CPU the backbone runs `fp32` by default** (maximum fidelity). Need more speed? Pass `Vieneu(precision="int8")` — ~1.6× faster and ~4× smaller, but it requires a CPU with VNNI (AVX-512 VNNI / AVX-VNNI); on older CPUs int8 can produce garbled audio. `precision` only affects the CPU/ONNX path; on GPU it's ignored (PyTorch). > > 🪶 **Still too slow, or deploying on a phone / ARM board?** Try **[VieNeu-TTS v3 Nano (preview)](#v3-nano)** — `Vieneu(mode="v3nano")`, ~3× faster than Turbo fp32 on CPU (RTF 0.11–0.22 on a desktop CPU), but **noticeably lower quality** (especially English / bilingual), 24 kHz, 11 preset voices + voice cloning. Details and caveats in the [v3 Nano section](#v3-nano) below. ```bash pip install vieneu ``` **GPU (CUDA)** — only if you have an NVIDIA GPU. On Linux `pip install "vieneu[cuda]"` is enough (PyPI torch ships CUDA there); on Windows install the CUDA torch **first** as below. > ℹ️ **How fast is the GPU path?** Since 3.7.0 every audio frame is **one CUDA > graph** (acoustic decoder + sampling + repetition penalty + backbone step in a > single replay — no `torch.compile`, no C++ toolchain needed). Measured on an > RTX 3060: a 3.5 s sentence in **0.36 s**; a 2-chunk paragraph (19 s) in > **1.4 s**; 16 chunks (154 s) in **2.8 s** (RTF 0.02) — previously 2.3 s / > 8.7 s / 16.7 s. The first call for each batch size pays ~0.5 s to capture the > graph (kept afterwards; servers can call `warm_fused()` at start-up). > `VIENEU_FUSED_FRAME=0` restores the plain loop. ```bash pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128 pip install "transformers==4.57.6" # pinned — most stable transformers for the GPU SDK pip install vieneu ``` ```python import time from vieneu import Vieneu # Default = v3 Turbo (48 kHz). GPU → PyTorch (auto-detected). vieneu = Vieneu() # On a GPU machine you can still switch to ONNX/CPU if you prefer: Vieneu(backend="onnx") # 1. Built-in voice by name — no reference clip needed print("🔊 Generating speech...") start_time = time.time() audio = vieneu.infer("[cười] Trời ơi, cái giọng nó tự nhiên mà nó mượt mà dã man, nghe không khác gì người thật luôn. Giờ thì tha hồ mà quẩy content với cả kho giọng nói đa dạng, đủ mọi sắc thái biểu cảm. Mọi người bật loa lên rồi cùng trải nghiệm thử với mình nhé!", voice="Phạm Tuyên") elapsed_time = time.time() - start_time vieneu.save(audio, "output.wav") print("✅ Saved to output.wav") # Tính RTF (Real-Time Factor) sample_rate = 48000 audio_duration = len(audio) / sample_rate rtf = elapsed_time / audio_duration print(f"\n⏱️ Thời gian xử lý: {elapsed_time:.3f}s") print(f"🎵 Thời lượng audio: {audio_duration:.3f}s") print(f"📊 RTF: {rtf:.4f} ({'nhanh hơn' if rtf < 1 else 'chậm hơn'} real-time {1/rtf:.2f}x)" if rtf > 0 else "") # List the built-in voices voices = vieneu.list_preset_voices() print(f"\n🎙️ {len(voices)} built-in voices available:") for label, voice_id in voices: print(f" - {label} ({voice_id})") # 2. ⚡ Batch on GPU: infer_batch() runs many texts in ONE batched forward — same API. # On a CUDA GPU the chunks from every text share each forward step (big throughput # win). On CPU it still WORKS (no error) — just sequentially, so there's no batch # gain. Batch caps at max_batch_size (default 32; tune via Vieneu(max_batch_size=64) # or infer_batch(..., batch_size=64), or batch_size=1 to disable). A single long # infer() also auto-batches its own chunks. For real-time use, infer_stream() is the # streaming twin (GPU: 16 concurrent streams — see "Streaming" below). Uncomment to # try (GPU recommended): # # import time # texts = [ # "Chào cả nhà, hôm nay mình sẽ hướng dẫn các bạn cách cài đặt và sử dụng bộ giọng đọc mới.", # "Giọng nghe cực kỳ tự nhiên và truyền cảm, lại có thể chuyển đổi biểu cảm một cách linh hoạt.", # "Nếu thấy hữu ích, các bạn nhớ để lại một lượt thích và chia sẻ video này cho mọi người nhé!", # ] * 10 # 30 texts — enough to fill the batch and really show the GPU throughput win # t0 = time.time() # audios = vieneu.infer_batch(texts, voice="Minh Quân Pro") # elapsed = time.time() - t0 # total_audio = sum(len(a) for a in audios) / 48_000 # print(f"⚡ {len(texts)} texts | audio {total_audio:.1f}s | wall {elapsed:.1f}s | RTF {elapsed/total_audio:.3f}") # for i, a in enumerate(audios): # vieneu.save(a, f"batch_{i}.wav") ``` ### Streaming (real-time) 🔊 > v3 Turbo streams **frame by frame** on both backends. **GPU** (PyTorch): first audio in **~115 ms** and **16 concurrent streams** on one RTX 3060 (continuous batching — one CUDA graph serves every `infer_stream` call, each keeping RTF ≈ 0.5–0.6). **CPU** (ONNX): first audio in ~140 ms (int8) / ~300 ms (fp32), one stream (two with int8). Just iterate `infer_stream`: ```python from vieneu import Vieneu vieneu = Vieneu() # GPU → PyTorch + stream scheduler; no GPU → ONNX/CPU for chunk in vieneu.infer_stream("Xin chào các bạn!", voice="Mai Anh"): play(chunk) # np.float32 @ 48 kHz — play/write as it arrives ``` Calling `infer_stream` from many threads at once is the intended way to serve many listeners on a GPU (`Vieneu(max_streams=16)` sets the ceiling). An **OpenAI-compatible streaming API** (`POST /v1/audio/speech`, `pcm`/`wav`, chunked or SSE — works with the OpenAI SDK, Pipecat, LiveKit, …) is in [`apps/openai_speech.py`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/apps/openai_speech.py): ```bash # Pick ONE of these — all serve http://localhost:8000/v1/audio/speech uv run python -m apps.openai_speech # from the repo (auto-detects GPU/CPU) docker compose -f docker/docker-compose.yml --profile api-gpu up # or: Docker, GPU docker compose -f docker/docker-compose.yml --profile api-cpu up # or: Docker, CPU only ``` 📊 **[docs/streaming.md](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/streaming.md)** — every measurement on an RTX 3060 (TTFA / RTF / streams vs `max_streams`), estimates for smaller GPUs, and the CPU numbers. The older browser demo is still at [`apps/web_stream.py`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/apps/web_stream.py). ### Available Voices The v3 Turbo engine includes **25 preset voices** covering **3 regions** (North, Central, South) with diverse genders and speaking characters. `list_preset_voices()` (and the Web UI / API voice lists) show them in this order: - ⭐ **Editors' picks** — the 10 we recommend starting with, hand-selected for naturalness and stability: **Adam bựa, Trúc Ly, Anh Khôi, Mai Anh, Minh Quân Pro** *(default; `"Minh Quân"` still works as an alias)*, **Thùy Dung, Thiền Tâm Đức, Ngọc Huyền, Quang Sơn, Ngọc Trân** - **Northern (Bắc)**: Minh Đức, Phạm Tuyên, Xuân Vĩnh, Thanh Bình, Ngọc Linh, Đoan Trang, Quỳnh Anh, Mạnh Dũng (+ picks above) - **Central (Trung)**: Quang Sơn, Ngọc Trân - **Southern (Nam)**: Adam, Thái Sơn, Thục Đoan, Minh Triết, Mỹ Duyên, Đức Trí, Kim Thanh (+ Thùy Dung) ## Reading style — **deprecated** ⚠️ :::warning **`style` is deprecated on v3 Turbo and has no effect.** The reading style is already baked into the reference itself (the speaker embedding + reference codes of the preset voice or of your cloned clip), so every generation follows the reference and comes out in its natural reading style. The `style` argument is **still accepted** by `infer`, `infer_stream`, `infer_batch` and `add_voice` so existing code keeps running — whatever you pass (`"tin_tuc"`, `"doc_truyen"`, …) is simply ignored. New code should just omit it. ::: ```python # Old code — still runs, but `style` is ignored audio = vieneu.infer("Bản tin sáng nay.", voice="Minh Quân Pro", style="tin_tuc") # New code — pick the reading character through the voice / reference clip instead audio = vieneu.infer("Bản tin sáng nay.", voice="Minh Quân Pro") ``` ## Emotion cues (experimental) Inline tags are supported anywhere in the text: `[cười]` (chuckle), `[thở dài]` (sigh), `[hắng giọng]` (clear throat). ```python audio = vieneu.infer("Nghe hay quá đi [cười]. Để mình nói tiếp [hắng giọng].", voice="Minh Quân Pro") ``` ## Voice cloning Clone any voice from a short reference clip. The clip is cleaned up automatically (background noise removed, and trimmed to ≤ 8 seconds) before cloning — keep `denoise=True` unless your clip is already clean. ```python audio = vieneu.infer( "Đây là giọng được nhân bản tức thì.", ref_audio="my_voice.wav", # a 3–8s reference clip denoise=True, # default; set False if the clip is already clean ) vieneu.save(audio, "cloned.wav") ``` ## Save & reuse a cloned voice Register a reference once with `add_voice`, then use it by name like a built-in voice. ```python # Enroll a voice (denoises + extracts the speaker profile once) vieneu.add_voice("Giọng của tôi", "my_voice.wav") # Now reuse it anywhere, including the conversation mode audio = vieneu.infer("Câu này dùng giọng đã lưu.", voice="Giọng của tôi") # Persist your voices so they load next session vieneu.save_voices() # writes to the default voices file # vieneu.remove_voice("Giọng của tôi") # Add a voice you already cleaned yourself → skip denoising vieneu.add_voice("Giọng sạch", "already_clean.wav", denoise=False) ``` ## Clean up a clip on its own Get the denoised audio without synthesizing anything (e.g. to inspect or store it): ```python wav, sr = vieneu.denoise("noisy.wav", out_path="clean.wav") # 44.1 kHz mono ``` > **Note:** `denoise`, `add_voice`, and voice cloning work on every backend — the > torch-free CPU/ONNX install included (the whole cloning pipeline runs on > onnxruntime + soxr + kaldi-native-fbank). **v3 Nano** below clones the same way (its cloning graphs are fetched on first use). ## v3 Nano (preview) — for edge devices / weak CPUs only 🪶 {#v3-nano} :::warning **v3 Turbo remains the default and the recommended model.** Use v3 Nano only when Turbo is too slow on your hardware (old laptops, mini PCs, ARM boards, CPUs without AVX-512/VNNI where the int8 Turbo build produces garbled audio). Nano is a 48M-parameter flow-matching model (ONNX, CPU, torch-free) and it **trades quality for speed**: - **Lower quality than v3 Turbo — most noticeably on English and code-switched (En-Vi) text.** Vietnamese is close; English words come out with a Vietnamese accent and are less stable. - **24 kHz** output (Turbo: 48 kHz). - **11 preset voices + voice cloning** (`ref_audio`, `add_voice`, `encode_reference` work like Turbo; the three cloning graphs, ~110 MB, download on first use). - **No frame-level streaming** — `infer_stream` yields one finished chunk at a time. ::: Measured on the same desktop CPU (12th-gen Intel i7, 6 ONNX Runtime threads, ~9 s of speech): | Engine | RTF ↓ | Sample rate | Load time | |---|---|---|---| | v3 Turbo ONNX fp32 (default on CPU) | 0.62 | 48 kHz | ~19 s | | v3 Turbo ONNX int8 | 0.37 | 48 kHz | ~14 s | | **v3 Nano, 16 steps, cfg 3** (default) | **0.22** | 24 kHz | ~3 s | | **v3 Nano, 8 steps, sway −1** | **0.11** | 24 kHz | ~3 s | RTF = compute time ÷ audio duration (lower is faster; 0.22 = 4.5× faster than real time). The ratio carries over to slower machines: expect Nano to be roughly **1.7× faster than Turbo int8** and **~3× faster than Turbo fp32**, with a 282 MB download instead of Turbo's. ```python from vieneu import Vieneu tts = Vieneu(mode="v3nano") # ONNX, CPU, torch-free audio = tts.infer("Xin chào, mình là giọng đọc của VieNeu Nano.", voice="Minh Quân") tts.save(audio, "nano.wav") # 24 kHz tts.list_preset_voices() # Adam, Ái Hân, Mỹ Duyên, Đức Trí, Hữu Quân, Xuân Tiên, Mai Anh, Trúc Ly, Anh Khôi, Minh Quân, Mạnh Dũng audio = tts.infer("Bản nhanh cho máy rất yếu.", voice="Ái Hân", steps=8, sway=-1) # ~2× faster ``` Knobs: `steps` (Euler steps, 16 default; 8 ≈ 2× faster, slightly rougher — pair with `sway=-1`), `cfg` (classifier-free guidance, 3.0 default; `cfg=0` halves compute but hurts intelligibility), `speed`, `seed`, `threads`. Emotion cues `[cười]` `[thở dài]` `[hắng giọng]` work as on Turbo. --- # Install & backends Source: https://docs.vieneu.io/docs/sdk/standard-mode One package, two engines. `Vieneu()` picks the engine from your hardware; the API is the same on both. | You have | Engine | Install | Notes | |---|---|---|---| | CPU only, macOS | ONNX Runtime, **torch-free** | `pip install vieneu` | 48 kHz v3 Turbo, cloning and emotion cues included. PyTorch is never installed. | | NVIDIA GPU | PyTorch (CUDA) | `pip install "vieneu[cuda]"` | Batched automatically; every frame is one CUDA graph since 3.7.0. | | Weak CPU / ARM board | ONNX, **v3 Nano** | `pip install vieneu` + `Vieneu(mode="v3nano")` | ~3× faster than Turbo fp32, 24 kHz, lower quality (esp. English). | ## CPU (default) ```bash pip install vieneu ``` ```python from vieneu import Vieneu tts = Vieneu() # v3 Turbo, ONNX, fp32 audio = tts.infer("Xin chào bạn", voice="Minh Quân Pro") tts.save(audio, "output.wav") # 48 kHz WAV ``` - **Precision.** `fp32` is the default for maximum fidelity. `Vieneu(precision="int8")` is ~1.6× faster and ~4× smaller, but needs a CPU with VNNI (AVX-512 VNNI / AVX-VNNI). On older CPUs int8 can produce garbled audio; if that happens, go back to fp32 or try v3 Nano. - **Fastest CPU install.** From a checkout of the repo, `uv sync` reproduces the locked environment that pins the optimised ONNX Runtime build. It is measurably faster than a plain `pip install`. - **Apple Silicon.** Use the CPU/ONNX path. It is faster than the MPS/PyTorch build for v3 Turbo. ## GPU (CUDA) On Linux the PyPI torch wheel already ships CUDA: ```bash pip install "vieneu[cuda]" ``` On Windows install the CUDA torch **first**, then a pinned transformers, then vieneu: ```bash pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128 pip install "transformers==4.57.6" pip install vieneu ``` ```python tts = Vieneu() # CUDA detected → PyTorch engine tts = Vieneu(backend="onnx") # force the CPU engine on a GPU machine ``` `precision` only applies to the ONNX path and is ignored on GPU. Throughput numbers and batching are on the [GPU batching](/docs/sdk/fast-mode) page. ## v3 Nano (preview) A 48M-parameter flow-matching model for hardware where Turbo is too slow. Same cloning API, 11 preset voices, 24 kHz output, no frame-level streaming (chunks arrive whole). ```python tts = Vieneu(mode="v3nano") audio = tts.infer("Bản nhanh cho máy rất yếu.", voice="Ái Hân", steps=8, sway=-1) ``` Measured on a 12th-gen i7 desktop, 6 ONNX threads (RTF = compute ÷ audio duration, lower is faster): | Engine | RTF | Sample rate | Load | |---|---|---|---| | v3 Turbo fp32 (CPU default) | 0.62 | 48 kHz | ~19 s | | v3 Turbo int8 | 0.37 | 48 kHz | ~14 s | | v3 Nano, 16 steps (default) | 0.22 | 24 kHz | ~3 s | | v3 Nano, 8 steps, sway −1 | 0.11 | 24 kHz | ~3 s | Knobs: `steps` (16 default, 8 ≈ 2× faster), `cfg` (3.0 default, 0 halves compute but hurts intelligibility), `speed`, `seed`, `threads`. ## Models and cache Backbone and codec weights download from Hugging Face on first use and are cached under `~/.cache/huggingface/hub/`. Turbo's cloning pipeline runs on onnxruntime + soxr + kaldi-native-fbank; Nano fetches its three cloning graphs (~110 MB) the first time you clone. ## Legacy backends (v1 / v2) The GGUF (llama-cpp) and LMDeploy backends serve **VieNeu-TTS v1/v2** only and are no longer updated. They live behind `uv sync --group gpu` in the repo and `pip install "vieneu[legacy]"`. New projects should stay on v3 Turbo. --- # GPU batching Source: https://docs.vieneu.io/docs/sdk/fast-mode On a CUDA GPU v3 Turbo batches automatically. A long `infer()` batches its own chunks; `infer_batch()` batches many texts in one forward pass. Same API, no code change. :::note Renamed page This page used to describe the LMDeploy "fast mode" for VieNeu-TTS v2. That backend is legacy now; see the bottom of the page. ::: ## infer_batch ```python from vieneu import Vieneu tts = Vieneu() # CUDA → PyTorch engine texts = [ "Chào cả nhà, hôm nay mình sẽ hướng dẫn các bạn cách cài đặt bộ giọng đọc mới.", "Giọng nghe cực kỳ tự nhiên và truyền cảm.", "Nếu thấy hữu ích, nhớ để lại một lượt thích nhé!", ] * 10 audios = tts.infer_batch(texts, voice="Minh Quân Pro") for i, a in enumerate(audios): tts.save(a, f"batch_{i}.wav") ``` - Chunks from every text share each forward step, which is where the throughput win comes from. - Batch size caps at `max_batch_size` (default 32). Tune with `Vieneu(max_batch_size=64)` or `infer_batch(..., batch_size=64)`; `batch_size=1` disables batching. - On CPU `infer_batch` still works, just sequentially, so there is no gain. - For real-time playback use [`infer_stream`](/docs/sdk/streaming) instead. It is the streaming twin and serves 16 concurrent listeners on one GPU. ## CUDA graphs (3.7.0+) Every audio frame is a single CUDA graph replay: acoustic decoder, sampling, repetition penalty and the backbone step in one shot. No `torch.compile`, no C++ toolchain. Measured on an RTX 3060: | Input | Audio | Wall time | Before 3.7.0 | |---|---|---|---| | One sentence | 3.5 s | 0.36 s | 2.3 s | | Two-chunk paragraph | 19 s | 1.4 s | 8.7 s | | 16 chunks | 154 s | 2.8 s (RTF 0.02) | 16.7 s | The first call for each batch size pays ~0.5 s to capture the graph, then keeps it. Servers can call `tts.warm_fused()` at start-up. `VIENEU_FUSED_FRAME=0` restores the plain loop if you need to debug. ## Requirements - NVIDIA GPU. An RTX 3060 (12 GB) gives the numbers above; ~6 GB is enough for inference. - Install per [Install & backends](/docs/sdk/standard-mode#gpu-cuda): CUDA torch 2.8.0 first on Windows, `transformers==4.57.6`, then `vieneu`. ## Fine-tuned models A LoRA-merged v3 Turbo keeps the full API, batching included: ```python tts = Vieneu(mode="v3turbo", backbone_repo="finetune/output/my_voice/merged") ``` See [Fine-tuning](/docs/advanced/fine-tuning) and the repo's [`finetune/README.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/finetune/README.md). ## Legacy: LMDeploy "fast" mode (v2 only) `Vieneu(mode="fast")` loads VieNeu-TTS v1/v2 through LMDeploy. It is kept for existing deployments (`pip install "vieneu[legacy]"`, or `uv sync --group gpu` in the repo) and receives no updates. v3 Turbo on PyTorch is faster and needs no extra runtime. --- # OpenAI-compatible server Source: https://docs.vieneu.io/docs/sdk/remote-mode `apps/openai_speech.py` in the repo serves `POST /v1/audio/speech` exactly like OpenAI's TTS endpoint (`pcm`/`wav`, chunked body or SSE). The OpenAI SDK, Pipecat, LiveKit Agents and the Vercel AI SDK work by changing `base_url`. ## Start the server Pick one; all listen on `http://localhost:8000`: ```bash uv run python -m apps.openai_speech # from a repo checkout, auto-detects GPU/CPU docker compose -f docker/docker-compose.yml --profile api-gpu up # Docker, GPU docker compose -f docker/docker-compose.yml --profile api-cpu up # Docker, CPU only (torch-free) ``` Measure time-to-first-audio and RTF on your own machine: ```bash uv run python examples/openai_speech_client.py --bench 8 ``` ## Call it ```python from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="x") with client.audio.speech.with_streaming_response.create( model="vieneu-v3-turbo", voice="Mai Anh", input="Xin chào! Đây là chế độ streaming của VieNeu.", response_format="pcm", ) as r: for chunk in r.iter_bytes(4096): # s16le 48 kHz mono, as it is generated play(chunk) ``` ## Endpoints | Method | Path | Purpose | |---|---|---| | `POST` | `/v1/audio/speech` | Synthesize; `response_format` `pcm` or `wav`, streamed | | `GET` | `/v1/models` | Model list | | `GET` | `/v1/voices` | Preset and enrolled voices | | `POST` | `/v1/voices` | Clone from an uploaded clip | | `GET` | `/health` | Liveness | Concurrency is capped per backend with `VIENEU_MAX_STREAMS` (default 16 on GPU, 1 on CPU). Requests beyond the cap wait in a small queue, then get `429`. First audio arrives in ~115 ms with 16 streams on an RTX 3060; ~140–300 ms and 1–2 streams on CPU. Full numbers: [`docs/streaming.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/streaming.md). ## Web UI in Docker ```bash docker compose -f docker/docker-compose.yml --profile gpu up # or --profile cpu → http://localhost:7860 ``` Production images and builds: the repo's [`docs/Deploy.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/Deploy.md) and our [Docker](/docs/deployment/docker) page. ## Hosted instead of self-hosted If you do not want to run a GPU, the [VieNeu Cloud API](/docs/cloud-api/openai-compatible) exposes the same OpenAI-compatible shape at `api.vieneu.io`, including the v4 engine. ## Legacy: v2 `remote` mode (deprecated) :::warning The LMDeploy server on port 23333 and `Vieneu(mode="remote")` only work with **VieNeu-TTS v2**, which is no longer updated. They are kept for existing deployments. For v3 Turbo use the streaming server above. ::: ```bash docker run --gpus all -p 23333:23333 -v huggingface_cache:/root/.cache/huggingface pnnbao/vieneu-tts:latest --tunnel pip install "vieneu[legacy]" ``` ```python from vieneu import Vieneu tts = Vieneu(mode="remote", api_base="http://your-server-ip:23333/v1", model_name="pnnbao-ump/VieNeu-TTS-v2", emotion="natural") audio = tts.infer(text="Chào bạn!") tts.save(audio, "remote_output.wav") ``` Fine-tuned v3 Turbo models are not served by that container; load them with the SDK (`Vieneu(mode="v3turbo", backbone_repo=...)`). See [Remote server](/docs/deployment/remote-server) for the old Docker flags. --- # Streaming Source: https://docs.vieneu.io/docs/sdk/streaming v3 Turbo streams **frame by frame** on both engines. Iterate `infer_stream` and play or write each chunk as it arrives. ```python from vieneu import Vieneu tts = Vieneu() # GPU → PyTorch + stream scheduler; CPU → ONNX for chunk in tts.infer_stream("Xin chào các bạn!", voice="Mai Anh"): play(chunk) # np.float32 @ 48 kHz ``` ## Latency and concurrency | Engine | First audio | Concurrent streams | |---|---|---| | GPU (PyTorch), RTX 3060 | ~115 ms | 16 (32 max), each at RTF ≈ 0.5–0.6 | | CPU (ONNX) int8 | ~140 ms | 2 | | CPU (ONNX) fp32 | ~300 ms | 1 | On a GPU the streams share one CUDA graph through continuous batching. Calling `infer_stream` from many threads at once is the intended way to serve many listeners; `Vieneu(max_streams=16)` sets the ceiling. The first request after the GPU has idled pays an extra 100–300 ms until it clocks up. Every measurement (TTFA and RTF against `max_streams`, estimates for smaller GPUs, CPU numbers) is in the repo's [`docs/streaming.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/streaming.md). ## Chunk format Each chunk is a `numpy.float32` array at 48 kHz, mono. Convert to 16-bit PCM for most audio devices: ```python import numpy as np pcm16 = (np.clip(chunk, -1, 1) * 32767).astype(np.int16).tobytes() ``` ## Serving over HTTP To expose streaming to other processes or languages, run the OpenAI-compatible server from the repo. It streams `pcm`/`wav` as chunked body or SSE from `POST /v1/audio/speech`, so the OpenAI SDK, Pipecat and LiveKit work by changing `base_url`. See [OpenAI-compatible server](/docs/sdk/remote-mode). ## v3 Nano Nano has no frame-level streaming: `infer_stream` yields one finished chunk at a time. Use Turbo when time-to-first-audio matters. ## Hosted alternative The same streaming shape is available without a GPU from the [Cloud API streaming endpoint](/docs/cloud-api/streaming). --- # Voice cloning Source: https://docs.vieneu.io/docs/sdk/voice-cloning Clone any voice from a **3–8 second** clip. No transcript, no fine-tuning. The clip is denoised and trimmed to ≤ 8 s automatically before cloning. ```python from vieneu import Vieneu tts = Vieneu() audio = tts.infer( "Đây là giọng được nhân bản tức thì.", ref_audio="my_voice.wav", # 3–8 s reference clip denoise=True, # default; False if the clip is already clean ) tts.save(audio, "cloned.wav") ``` :::info No `ref_text` on v3 v1/v2 needed the exact transcript of the reference clip (`ref_text`). v3 Turbo extracts a speaker embedding and reference codes instead, so a transcript is not required. ::: ## Enrol once, reuse by name ```python tts.add_voice("Giọng của tôi", "my_voice.wav") # denoise + extract the profile once audio = tts.infer("Câu này dùng giọng đã lưu.", voice="Giọng của tôi") tts.save_voices() # persist to the default voices file # tts.remove_voice("Giọng của tôi") tts.add_voice("Giọng sạch", "already_clean.wav", denoise=False) ``` Enrolled voices work everywhere a preset does, including `infer_batch`, `infer_stream` and the conversation mode. ## Denoise on its own ```python wav, sr = tts.denoise("noisy.wav", out_path="clean.wav") # 44.1 kHz mono ``` ## Reading style follows the reference The `style` argument (`tin_tuc`, `doc_truyen`, …) is **deprecated and ignored** on v3 Turbo. The reading character is baked into the reference: pick a preset voice or a clip that already reads the way you want. Passing `style` still runs for backwards compatibility. ## Emotion cues (experimental) Inline tags work with cloned voices too: `[cười]` (chuckle), `[thở dài]` (sigh), `[hắng giọng]` (clear throat). ```python audio = tts.infer("Nghe hay quá đi [cười]. Để mình nói tiếp [hắng giọng].", voice="Giọng của tôi") ``` ## Backends `denoise`, `add_voice` and cloning work on every backend, including the torch-free CPU/ONNX install (the pipeline runs on onnxruntime + soxr + kaldi-native-fbank). **v3 Nano** clones the same way; its three cloning graphs (~110 MB) download on first use. ## Tips for a good reference - 3–8 s of one speaker, no music, no second voice. - Natural, continuous speech beats isolated words. - Keep `denoise=True` unless you cleaned the clip yourself. - Want a tighter match or a specific reading style? [Fine-tune with LoRA](/docs/advanced/fine-tuning) on 10–30 minutes of audio. ## Higher fidelity: v4 on the Cloud API The proprietary **VieNeu-TTS v4** reproduces a reference with near-original speaker similarity. It is not open source and is available only through the [Cloud API](/docs/cloud-api/overview). --- # Vieneu() Factory Source: https://docs.vieneu.io/docs/api/factory The main entry point for creating a VieNeu-TTS instance. ## Signature ```python Vieneu(mode="standard", **kwargs) ``` ## Parameters | Parameter | Type | Default | Description | |-----------|------|---------|-------------| | `mode` | `str` | `"standard"` | Backend mode: `"standard"`, `"fast"`, `"remote"`, `"xpu"` | ### Standard Mode kwargs | Parameter | Type | Default | Description | |-----------|------|---------|-------------| | `backbone_repo` | `str` | `"pnnbao-ump/VieNeu-TTS-0.3B-q4-gguf"` | HuggingFace repo or local path | | `backbone_device` | `str` | `"cpu"` | `"cpu"`, `"cuda"`, `"mps"` | | `codec_repo` | `str` | `"neuphonic/distill-neucodec"` | Codec model repo | | `codec_device` | `str` | `"cpu"` | Device for codec | | `hf_token` | `str` | `None` | HuggingFace token for private models | ### Remote Mode kwargs | Parameter | Type | Default | Description | |-----------|------|---------|-------------| | `api_base` | `str` | required | Server URL (e.g., `"http://host:23333/v1"`) | | `model_name` | `str` | required | Model ID on the server | ## Returns An instance of `BaseVieneuTTS` (subclass depends on `mode`). ## Examples ```python # Default: GGUF Q4 on CPU tts = Vieneu() # PyTorch on GPU tts = Vieneu(backbone_repo="pnnbao-ump/VieNeu-TTS-0.3B", backbone_device="cuda") # LMDeploy fast mode tts = Vieneu(mode="fast", backbone_repo="pnnbao-ump/VieNeu-TTS") # Remote client tts = Vieneu(mode="remote", api_base="http://server:23333/v1", model_name="pnnbao-ump/VieNeu-TTS") ``` --- # Inference Methods Source: https://docs.vieneu.io/docs/api/infer ## `infer()` ```python audio = tts.infer( text: str, ref_audio: str = None, ref_codes: Tensor = None, ref_text: str = None, voice: dict = None, max_chars: int = 256, silence_p: float = 0.15, crossfade_p: float = 0.0, temperature: float = 1.0, top_k: int = 50, skip_normalize: bool = False, ) ``` ### Parameters | Parameter | Type | Description | |-----------|------|-------------| | `text` | `str` | Text to synthesize | | `ref_audio` | `str` | Path to reference audio for voice cloning | | `ref_codes` | `Tensor` | Pre-encoded reference codes | | `ref_text` | `str` | Transcript of reference audio | | `voice` | `dict` | Preset voice dict from `get_preset_voice()` | | `max_chars` | `int` | Max characters per chunk (default 256) | | `silence_p` | `float` | Silence duration between chunks in seconds | | `crossfade_p` | `float` | Crossfade duration between chunks | | `temperature` | `float` | Sampling temperature | | `top_k` | `int` | Top-k sampling | | `skip_normalize` | `bool` | Skip text normalization | ### Returns `numpy.ndarray` — Audio waveform at 24 kHz. ### Voice Priority 1. `voice` dict (from preset) 2. `ref_audio` + `ref_text` 3. `ref_codes` + `ref_text` 4. Default preset voice --- ## `infer_batch()` ```python audios = tts.infer_batch(texts: List[str], ...) ``` Returns `List[numpy.ndarray]`. PyTorch mode uses true batch generation; GGUF processes sequentially. --- ## `infer_stream()` ```python for chunk in tts.infer_stream(text: str, ...): play_audio(chunk) ``` Yields `numpy.ndarray` chunks (GGUF only). --- ## `save()` ```python tts.save(audio: numpy.ndarray, output_path: str) ``` --- ## `encode_reference()` ```python codes = tts.encode_reference(ref_audio_path: str) # Returns: torch.Tensor ``` --- ## `close()` ```python tts.close() # Or use context manager: with Vieneu() as tts: audio = tts.infer(text="...") ``` --- # Voice Management Source: https://docs.vieneu.io/docs/api/voice-management ## Preset Voices ### `list_preset_voices()` ```python voices = tts.list_preset_voices() # Returns: List[tuple[str, str]] → [(description, voice_id), ...] ``` ### `get_preset_voice()` ```python voice = tts.get_preset_voice(voice_name: str = None) # Returns: dict → {"codes": Tensor, "text": str} ``` ### Using a Preset Voice ```python voices = tts.list_preset_voices() voice = tts.get_preset_voice("bac_si_tuyen") audio = tts.infer(text="Chào bạn!", voice=voice) ``` ## LoRA Adapters ### `load_lora_adapter()` ```python success = tts.load_lora_adapter( lora_repo_id: str, hf_token: str = None, ) ``` ### `unload_lora_adapter()` ```python success = tts.unload_lora_adapter() ``` ## voices.json Format ```json { "default_voice": "voice_name", "presets": { "voice_name": { "description": "Description of the voice", "text": "Transcript of the reference audio", "codes": [42, 17, 89, 55, ...] } } } ``` --- # Docker Deployment Source: https://docs.vieneu.io/docs/deployment/docker ## Requirements - Docker - [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html) ## Docker Compose ```bash docker compose -f docker/docker-compose.yml --profile gpu up ``` ## Custom Docker Run ```bash docker run --gpus all -p 7860:7860 pnnbao/vieneu-tts:latest ``` Access the Web UI at `http://localhost:7860`. :::note Docker deployment currently supports **GPU only**. For CPU, install from source. ::: --- # Remote Server Source: https://docs.vieneu.io/docs/deployment/remote-server Deploy VieNeu-TTS as a high-performance API server powered by LMDeploy. ## Quick Start ```bash docker run --gpus all -p 23333:23333 pnnbao/vieneu-tts:serve ``` ## With Public Tunnel ```bash docker run --gpus all -p 23333:23333 pnnbao/vieneu-tts:serve --tunnel ``` ## Run 0.3B Model (Faster) ```bash docker run --gpus all pnnbao/vieneu-tts:serve --model pnnbao-ump/VieNeu-TTS-0.3B --tunnel ``` ## Serve a Fine-tuned Model ```bash docker run --gpus all \ -v $(pwd)/finetune/output:/workspace/models \ pnnbao/vieneu-tts:serve \ --model /workspace/models/merged_model --tunnel ``` ## Connecting Clients See [Remote Mode](/docs/sdk/remote-mode) for client SDK usage. --- # Custom Models Source: https://docs.vieneu.io/docs/advanced/custom-models ## Via SDK ```python from vieneu import Vieneu # Load from HuggingFace tts = Vieneu(backbone_repo="your-username/your-model") # Load from local path tts = Vieneu(backbone_repo="/path/to/your/model") ``` ## Via Web UI The Web UI provides a model selector where you can enter any HuggingFace repo ID or local path. ## LoRA Adapters ```python tts = Vieneu( backbone_repo="pnnbao-ump/VieNeu-TTS", backbone_device="cuda", ) tts.load_lora_adapter("your-username/your-lora-adapter") audio = tts.infer(text="Chào bạn!") tts.unload_lora_adapter() ``` :::note LoRA adapters require PyTorch backbone. Not supported with GGUF models. ::: --- # Fine-tuning Source: https://docs.vieneu.io/docs/advanced/fine-tuning Train VieNeu-TTS on your own voice or custom datasets using LoRA. ## Quick Start ```bash cd finetune python train.py --config config.yaml ``` ## Google Colab Use the provided notebook: `finetune/finetune_VieNeu-TTS.ipynb` ## Workflow 1. **Prepare data** — Audio files + transcripts 2. **Configure** — Edit training config (LoRA rank, learning rate, etc.) 3. **Train** — Run `train.py` 4. **Merge** — Merge LoRA weights into base model (optional) 5. **Use** — Load via `load_lora_adapter()` or serve merged model ## Documentation See the detailed guide at [`finetune/README.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/finetune/README.md). --- # Model Overview Source: https://docs.vieneu.io/docs/advanced/model-overview ## Available Models | Model | Format | Device | Quality | Speed | |-------|--------|--------|---------|-------| | VieNeu-TTS (0.5B) | PyTorch | GPU/CPU | Best | Very Fast (LMDeploy) | | VieNeu-TTS-0.3B | PyTorch | GPU/CPU | Great | Ultra Fast (2x) | | VieNeu-TTS Q8 GGUF | GGUF | CPU/GPU | Great | Fast | | VieNeu-TTS Q4 GGUF | GGUF | CPU/GPU | Good | Very Fast | | VieNeu-TTS-0.3B Q8 GGUF | GGUF | CPU/GPU | Great | Ultra Fast (1.5x) | | VieNeu-TTS-0.3B Q4 GGUF | GGUF | CPU/GPU | Good | Extreme (2x) | ## Architecture - **0.5B** — Fine-tuned from NeuTTS Air architecture. Maximum stability and quality. - **0.3B** — Trained from scratch on VieNeu-TTS-1000h. 2x faster, ultra-low latency. ## Technical Details | Spec | Value | |------|-------| | Training data | VieNeu-TTS-1000h (443,641 samples) | | Audio codec | NeuCodec | | Context window | 2,048 tokens | | Output sample rate | 24 kHz | | Watermark | Enabled by default | ## HuggingFace Links - [VieNeu-TTS (0.5B)](https://huggingface.co/pnnbao-ump/VieNeu-TTS) - [VieNeu-TTS-0.3B](https://huggingface.co/pnnbao-ump/VieNeu-TTS-0.3B) - [Training Dataset](https://huggingface.co/datasets/pnnbao-ump/VieNeu-TTS-1000h)