# VieNeu > Vietnamese text-to-speech with instant voice cloning, available both as a > hosted HTTP API and as a package that runs the model on your own machine. VieNeu ships as two separate products that share a name and a voice catalogue, and nothing else — different install, different billing, different code: - **Cloud API** (`https://api.vieneu.io/api/v1`) — send text over HTTPS, get audio back. Billed per submitted character. Authenticated with an API key (`Authorization: Bearer vn_sk_…` or `X-API-Key`). This is what almost every integration wants. - **On-device SDK** — a Python package that runs the model locally. No network, no per-character cost; the hardware and setup are yours. ## Fastest way in: the hosted MCP server (no code) `https://api.vieneu.io/mcp` — a remote Model Context Protocol server (Streamable HTTP). Add it to Claude, ChatGPT, Cursor, VS Code, Claude Code or any MCP client and the assistant can find voices and synthesize Vietnamese speech for the user. - Auth: the user signs in with their VieNeu account through OAuth 2.1 (discovery at `/.well-known/oauth-protected-resource`; Client ID Metadata Documents or Dynamic Client Registration; PKCE S256), or the client sends a VieNeu API key as `X-API-Key` / `Authorization: Bearer`. Needs an account with an active token plan. - Tools: `list_capabilities` (what the connection can do, with examples), `list_voices`, `text_to_speech` (MP3 link for up to 1,500 characters, WAV job above; billed like the API), `get_speech_status`, `list_emotion_tags`, `get_token_balance`. In MCP Apps hosts the results render an inline audio player. - Audiobooks: `create_audiobook`, `add_audiobook_chapter` (the user's text as a script: narration, and each line of dialogue with its speaker), `start_audiobook` (billed per chapter as it renders), `get_audiobook`, `list_audiobooks`, `cancel_audiobook`. Books render in the background while the GPUs are idle (hours); one mastered MP3 per chapter. - Prompts: `what_can_vieneu_do`, `read_text`, `make_audiobook`, `find_voice`, `check_balance`. - Claude Code: `claude mcp add --transport http vieneu https://api.vieneu.io/mcp` - Cursor / VS Code `mcp.json`: `{"url": "https://api.vieneu.io/mcp"}` (VS Code: add `"type": "http"` under `servers`). - Guides: [overview](https://docs.vieneu.io/docs/integrations/mcp/index.md), [Claude](https://docs.vieneu.io/docs/integrations/mcp/claude-ai.md), [ChatGPT](https://docs.vieneu.io/docs/integrations/mcp/chatgpt.md), [Claude Code](https://docs.vieneu.io/docs/integrations/mcp/claude-code.md), [Cursor](https://docs.vieneu.io/docs/integrations/mcp/cursor.md), [VS Code](https://docs.vieneu.io/docs/integrations/mcp/vscode.md), [OpenAI API](https://docs.vieneu.io/docs/integrations/mcp/openai-api.md), [audiobooks](https://docs.vieneu.io/docs/integrations/mcp/audiobooks.md), [other clients](https://docs.vieneu.io/docs/integrations/mcp/other-clients.md), [troubleshooting](https://docs.vieneu.io/docs/integrations/mcp/troubleshooting.md). ## Machine-readable - [OpenAPI specification](https://docs.vieneu.io/openapi.json): the complete Cloud API contract, generated from the server's own route definitions. Every path, field, enum and response. Read this before reading anything below. - [Rendered reference](https://docs.vieneu.io/api-reference): the same document as a browsable page. ## Getting an API key Keys are created on the **Developer page** of the VieNeu web app: . The plaintext is shown once, at creation. - `vn_sk_…` — live keys. - `vn_test_…` — test keys. Identical in every respect (same billing, same rate limits, same endpoints) except that each request is capped at 100 whitespace-separated words. - A key is only issued against an **active token grant**. A new account has none, so key creation answers `403` `No active API token grant found.` until one exists; the Developer page offers a free 7-day trial that starts one, once per account. - At most 5 active keys per account. ## Cloud API essentials - Synchronous synthesis: `POST /v1/audio/speech`. This is a drop-in for OpenAI's endpoint of the same name, so an OpenAI SDK works unchanged against `base_url="https://api.vieneu.io/api/v1"`. Defaults to mp3; `wav`, `opus`, `pcm` and 8 kHz `ulaw` are available via `response_format`. - Asynchronous synthesis: `POST /v1/tts` returns a `jobId`; poll `GET /v1/tts/{jobId}`. - Streaming: `POST /v1/tts/stream` (length-prefixed frames, ending in a zero-length frame that proves the stream completed), or `stream_format` on `/v1/audio/speech` for plain chunked audio or OpenAI-shaped SSE. First audio lands in roughly 1–2 seconds on the standard plans; Enterprise runs on a dedicated node with a latency ceiling agreed per contract, down to 250 ms. Streaming carries `pcm` or `ulaw`: frames are encoded independently, so mp3 and opus only join cleanly in a non-streamed request. - Voices: `GET /v1/voices` (rich, no key needed) or `GET /v1/audio/voices` (ids only, for OpenAI-compatible clients — this one **does** need a key). A voice belongs to one engine — `v4` (the default) or `v3` — and the two catalogues are not interchangeable: most `v3` ids are opaque `vieneu-…` slugs that exist on no other engine, `v4`'s are Vietnamese display names, and a minority of names appear on both. Resolve ids at runtime with `GET /v1/voices?engine=…`; passing one from the wrong engine is a `400`. - Engines and capabilities: `GET /v1/engines` returns the live defaults, feature lists and billing multipliers, no key needed. `GET /v1/emotion-tags?engine=…` returns the reading styles and inline cue tags that engine can actually render — also no key. `emotion` is a **`v3`-only** field: `v4` has no styles and deletes any tag it cannot voice. - Voice agents: `POST /v1/vapi/speech` implements Vapi's `custom-voice` webhook (raw mono s16le PCM at the requested rate). Point a Vapi assistant at it, put the VieNeu API key in `server.secret`, and pick the voice with `?voiceId=`. - Voice cloning is **web-only** since 2026-09-23: a voice is cloned in the web Studio at https://vieneu.io/#/clone, not through the API. The old enrolment routes (`POST /v1/voices`, `DELETE /v1/voices/{id}`, `POST /v1/clone`, `POST /v1/upload`, `POST /v1/prepare`) answer `410` with `code: "CLONE_WEB_ONLY"` and charge nothing. A voice cloned there is listed by `GET /v1/voices` when you send your key (`"kind": "cloned"`) and its `clone_…` id works as `voiceId` on `POST /v1/tts` and `POST /v1/tts/stream`, at the same price as a catalogue voice — clones are `v4`, so they stream. - Also available: `POST /v1/dialogue` (multi-speaker in one call), `POST /v1/dub` (re-voice an existing recording), `POST /v1/srt` (dub a subtitle file into a timecode-aligned track). - Usage: `GET /v1/usage?from&to&scope=key|account` returns what the key (or the whole account) did over a window of up to 92 days — calls by outcome, tokens, characters, seconds of audio, generation time, per day / route / voice / error code. Read from the per-call ledger, so it matches what was billed exactly. - Webhooks instead of polling: `POST /v1/webhooks`, `GET /v1/webhooks`, `GET /v1/webhooks/{id}`, `DELETE /v1/webhooks/{id}`, `GET /v1/webhooks/{id}/deliveries`, `POST /v1/webhooks/{id}/rotate-secret`. Events are HMAC-SHA256 signed, at-least-once and unordered; dedupe on the event `id`. Retries fire at t+0, +0.5m, +2.0m, +5.5m, +13.0m and +28.5m, so size reconciliation thresholds above the longest gap of 15.5 minutes. - Billing is per submitted character × the engine multiplier, minimum 50 characters. `POST /v1/dub` is the one exception: it is billed by output audio duration (65 tokens/second, 150-token floor) and charged only on success. - **Rate limiting has two layers.** The application limiter counts your API key and is the one that matches what you bought — its refusals carry `X-RateLimit-Limit`, `X-RateLimit-Remaining`, `X-RateLimit-Reset` and `Retry-After`. In front of it, an nginx edge counts your **source IP** with both a request rate and a concurrent-connection cap, shared with everyone behind that address; its refusals are bare `429`/`503` responses carrying none of those headers. The presence or absence of `X-RateLimit-*` is how you tell which one you hit. `/v1/tts/stream` sits behind a tighter connection cap than the rest of `/api/v1`. - Every response carries `X-Request-Id`; error bodies repeat it as `traceId`. Quote it in support requests. - `aiRefine` defaults to **off** on the API: text is synthesized exactly as sent, with no AI moderation or normalization and no surcharge. Set it to `true` for the AI pass the web app uses. - The public API never returns `402`. An exhausted or expired grant is `403`; a daily or weekly cap is `429`. ## Docs - [Cloud API overview](https://docs.vieneu.io/docs/cloud-api/overview): base URL, auth, billing, all 32 operations, rate limits - [Quickstart](https://docs.vieneu.io/docs/cloud-api/quickstart): key, voice, one call — curl, Python and JavaScript - [OpenAI-compatible endpoint](https://docs.vieneu.io/docs/cloud-api/openai-compatible): request fields, formats, streaming, differences from OpenAI - [Streaming](https://docs.vieneu.io/docs/cloud-api/streaming): frame format, reference decoders in Python and JavaScript, latency - [Errors](https://docs.vieneu.io/docs/cloud-api/errors): both body shapes, every machine-readable code, the text validations - [Webhooks](https://docs.vieneu.io/docs/cloud-api/webhooks): event bodies, signature verification, delivery guarantees, retries - [Vapi](https://docs.vieneu.io/docs/cloud-api/vapi): the custom-voice webhook for voice agents - [Changelog](https://docs.vieneu.io/docs/cloud-api/changelog): API changes, with behaviour changes marked - [Integrations overview](https://docs.vieneu.io/docs/integrations/overview): which path to take for the tool you already use - [OpenAI-protocol clients](https://docs.vieneu.io/docs/integrations/openai-clients): Open WebUI, LiteLLM, LiveKit and friends — usually just a base_url - [MCP server](https://docs.vieneu.io/docs/integrations/mcp): the hosted MCP server — speech and audiobooks from Claude, ChatGPT, Cursor, VS Code and Claude Code - [n8n](https://docs.vieneu.io/docs/integrations/n8n): the community node, plus a zero-install HTTP Request path - [On-device SDK overview](https://docs.vieneu.io/docs/sdk/overview): the open-source vieneu package (v3 Turbo on CPU/GPU), mirrored monthly from the README - [Installation](https://docs.vieneu.io/docs/getting-started/installation): SDK setup and system requirements ## Notes - All text is Vietnamese-first; English words mixed into Vietnamese text are handled without markup. - Audio is 48 kHz by default, except `pcm` on `/v1/audio/speech`, which is 24 kHz to match what OpenAI documents. `opus` is always 48 kHz and `ulaw` always 8 kHz. - `pcm` and `ulaw` are headerless: read the rate from the `X-Sample-Rate` header.