Skip to main content

Drop-in for OpenAI-compatible apps

A lot of software already speaks POST /v1/audio/speech and lets you point it at a custom base URL. VieNeu implements that endpoint, so those apps can use Vietnamese voices with no plugin and no adapter.

This page is the setup sheet: what to put in which field, for each client.

The VieNeu side of every section below is exact. The client side is written as "which value goes in which kind of field", because setting labels move between versions — verify the field names against your build.

The three values​

Every client needs the same three things.

ValueWhat to enter
Base URLhttps://api.vieneu.io/api/v1 (see the trap below)
API keyyour VieNeu key — starts vn_sk_ (live) or vn_test_ (test)
Modeltts-1, tts-1-hd, gpt-4o-mini-tts, vieneu or vieneu-v4 → v4, the only engine since v3 was retired on 2026-09-24. Any other name is accepted and ignored — except a vieneu-… name that is not a live engine, which is a 400. vieneu-v3 is now one of those.

The one carve-out in that last row is the one a VieNeu user is most likely to trip. vieneu-v5 or vieneu-turbo does not fall through to the default engine the way whatever-tts does; it returns 400 invalid_request_error with param: "model", listing the engine names that do exist. The prefix is read as an explicit engine choice, and quietly rendering it on some other engine would bill at a rate you never picked. If you send both engine and model, engine wins.

vieneu-v3 is the case you are most likely to actually hit. It was a valid choice until v3 was retired from the cloud API on 2026-09-24, and now answers 400 with Model 'vieneu-v3' is retired. Engine "v3" was retired on 2026-09-24. Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues share no ids. Change the model to vieneu-v4 (or tts-1) and re-check the voice in the same pass — a vieneu-… slug saved from the old v3 catalogue will 400 on param: "voice" next.

A fourth field, voice, is optional but usually wanted: omitting it uses the first active voice on the resolved engine. What you cannot do is reuse OpenAI's names — alloy / nova / etc. are not mapped and return 400. Get real ids from GET /api/v1/audio/voices (below).

The base-URL trap​

The path is /api/v1/audio/speech. The doubled-looking /api/v1 is real — the server sets a global api prefix and mounts the public API at v1. So the value you type depends on what your client appends to it:

If the field means…Enter
"the OpenAI base — I append /audio/speech"https://api.vieneu.io/api/v1
"the host — I append /v1/audio/speech"https://api.vieneu.io/api
"the full endpoint URL"https://api.vieneu.io/api/v1/audio/speech

Most clients mean the first. A wrong pick is always a 404, and always JSON — but in VieNeu's platform error shape rather than the OpenAI envelope this endpoint otherwise uses:

{ "statusCode": 404, "message": "Cannot POST /api/v1/v1/audio/speech", "error": "Not Found" }

Two things to read off it:

  • The Cannot POST … message and the absence of the {"error":{"message","type","param","code"}} wrapper that every real /v1/audio/speech failure carries mean the URL shape is wrong, not the key. Do not go hunting your API key on this one.
  • The path inside the message is the URL your client actually built. Compare it to /api/v1/audio/speech and the difference tells you which row above you needed — the example here doubled /v1, so that client wanted https://api.vieneu.io/api.

There is no /v1/models​

VieNeu serves TTS only. There is no /v1/models route and no /v1/chat/completions. Two consequences:

  • A "Test connection" / "Verify" button that probes /models will report failure even though synthesis works. Ignore it and send a real request.
  • A model dropdown fed from /models will be empty. Type the model name by hand.

In Open WebUI, SillyTavern and LobeChat, this base URL belongs only in the audio/TTS provider setting — never in the general OpenAI/LLM endpoint setting, or the chat model breaks.

Prove it works first​

Before touching any client, confirm the base URL and the key with curl. Send $VIENEU_API_KEY — your own vn_sk_… or vn_test_… key — and no voice, so nothing but those two values can fail:

curl https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu." }' \
--output speech.mp3

Omitting voice makes the server pick the first active voice on the resolved engine, which is what keeps this step honest: a voice id you have not verified yet would 400 on param: "voice" and leave you unable to tell a bad voice from a bad key or a bad base URL. Pick a real voice in the next step, once this one has passed.

If that produces playable audio, everything after this is client configuration.

Getting voice ids​

# Which engine will my requests land on? No API key needed.
curl https://api.vieneu.io/api/v1/engines

# Voice ids for that engine. API key IS required here.
curl "https://api.vieneu.io/api/v1/audio/voices?engine=v4" \
-H "Authorization: Bearer $VIENEU_API_KEY"

GET /api/v1/engines returns one entry per enabled engine with key, isDefault, billingMultiplier and features. Use the one with isDefault: true — since v3 was retired on 2026-09-24 that is the only entry, v4.

Voice ids saved before 2026-09-24 may be dead

audio/voices now lists v4 only, filtered or not, and ?engine=v3 is a 400. But the retired V3 catalogue's ids were vieneu-… slugs, V4 ids are display names (Ngọc Lan), and the two shared almost no id space — so a voice a client saved from an unfiltered list before the retirement will 400 on param: "voice" now. Refresh the picker from the call above, and keep passing ?engine=v4: it costs nothing and keeps the request explicit.

Voice matching is case-insensitive but not diacritic-insensitive, and V4 ids contain spaces. A client that slugifies or strips accents before sending will 400 on every request.

Now re-run the smoke test with an id from that list, copied exactly as returned:

curl https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu.", "voice": "PASTE_AN_ID_HERE" }' \
--output speech.mp3

If the first curl worked and this one 400s on param: "voice", the id is the only thing that changed — check the engine filter and the accents before anything else.

Cloned voices do not work here

clone_… ids are rejected on /v1/audio/speech with a message naming the two routes that accept them: POST /v1/tts and POST /v1/tts/stream. No OpenAI-compatible client can reach a cloned voice.

Open WebUI​

Audio/TTS settings, OpenAI engine:

FieldValue
TTS engine / providerthe OpenAI-compatible option
API base URLhttps://api.vieneu.io/api/v1
API keyvn_sk_…
Modeltts-1
Voicea real id from audio/voices, e.g. Ngọc Lan

Notes:

  • Put this in the audio settings, not the model/connection settings.
  • The voice field must accept free text. If your build only offers OpenAI's six names, they will all 400 — verify on your build.
  • Response format: leave it at the default. Open WebUI plays mp3, which is also VieNeu's default.

SillyTavern​

TTS extension, OpenAI-compatible provider:

FieldValue
Providerthe OpenAI TTS option
API / base URLhttps://api.vieneu.io/api/v1
API keyvn_sk_…
Modeltts-1
Voice (per character)a real id from audio/voices

Notes:

  • SillyTavern assigns a voice per character. Every one of them must be a VieNeu id — a character left on alloy fails while the others work, which reads as a flaky integration.
  • If the extension only offers a fixed voice dropdown rather than a text field, it cannot address VieNeu voices. Verify on your build.
  • If your SillyTavern runs its TTS call from the browser rather than its own server, see Browser-side clients.

LobeChat​

TTS / audio settings, OpenAI provider:

FieldValue
OpenAI TTS base URL / proxy URLhttps://api.vieneu.io/api/v1
API keyvn_sk_…
Modeltts-1
Voicea real id from audio/voices

Notes:

  • LobeChat keeps separate settings for the chat provider and the TTS provider. This URL goes in the TTS one only.
  • LobeChat is commonly deployed so that the browser calls the TTS provider directly — see Browser-side clients before you debug anything else.

LiteLLM​

Add VieNeu as a model in config.yaml:

model_list:
- model_name: vieneu-tts
litellm_params:
model: openai/tts-1
api_base: https://api.vieneu.io/api/v1
api_key: os.environ/VIENEU_API_KEY

Then, with the proxy running, call it exactly as you would OpenAI:

curl http://localhost:4000/v1/audio/speech \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "vieneu-tts", "input": "Xin chào", "voice": "Ngọc Lan" }' \
--output speech.mp3

Notes:

  • LITELLM_MASTER_KEY is not your VieNeu key. It is LiteLLM's own proxy credential, whatever you set when you started the proxy — the port (:4000 above) is LiteLLM's too. Your vn_sk_… key appears only in config.yaml, reached through os.environ/VIENEU_API_KEY; the proxy is what attaches it to the upstream call.
  • The openai/ prefix on model tells LiteLLM to pass the request through in OpenAI's shape, which is what VieNeu answers. The name after it (tts-1) is what reaches VieNeu.
  • The VieNeu values above are exact; the surrounding LiteLLM keys are its standard model_list shape — confirm them against your LiteLLM version.
  • Do not enable any request-enrichment that adds body fields. VieNeu rejects unknown fields with 400 rather than ignoring them — see 400.

LiveKit Agents​

The OpenAI plugin takes a base URL and a key:

from livekit.plugins import openai

tts = openai.TTS(
model="vieneu-v4",
voice="Ngọc Lan",
base_url="https://api.vieneu.io/api/v1",
api_key="vn_sk_...",
)

Notes:

  • Pin vieneu-v4, or omit model. Either lands on v4, the only engine since v3 was retired on 2026-09-24. vieneu-v3 — which this endpoint used to accept with streaming, answering 200 and delivering v3's chunks in a burst at the end with no error telling you so — is now a 400 with param: "model" saying the engine was retired.
  • GET /v1/engines returns the live billing multiplier; see Engines.
  • For streaming, VieNeu needs both response_format: "pcm" and stream_format. response_format defaults to mp3, so setting only stream_format is a 400. Whether the plugin sends either, and whether it lets you add them, is the thing to verify on your build. Without them the call still works — it just returns a complete mp3 instead of a stream.
  • pcm is headerless and defaults to 24000 Hz on this route (the engines' native rate is 48000). Read the actual rate from the X-Sample-Rate response header and configure the pipeline to match, or the audio plays at the wrong pitch.

Pipecat​

import os
from pipecat.services.openai.tts import OpenAITTSService

tts = OpenAITTSService(
api_key=os.environ["VIENEU_API_KEY"],
base_url="https://api.vieneu.io/api/v1",
model="vieneu-v4",
voice="Ngọc Lan",
)

Notes:

  • The import path for Pipecat's OpenAI TTS service has moved between releases — check yours. The constructor arguments and the values above are what matter.
  • The same three streaming rules as LiveKit apply: pin vieneu-v4, send response_format: "pcm" and stream_format, and take the rate from X-Sample-Rate rather than assuming 48 kHz.
  • If your version hard-codes a response_format VieNeu does not accept (aac, flac), the request 400s naming the field. wav, mp3, opus, pcm and ulaw are the accepted values.

Browser-side clients​

In production, VieNeu's CORS policy allows only VieNeu's own origins. A client that calls /v1/audio/speech from the page rather than from its own server fails at the preflight, with no response body to explain it.

Whether a given LobeChat or SillyTavern deployment does its TTS server-side is a per-deployment question — verify on your build. If it is browser-side, put a proxy of your own in front (LiteLLM works well for this) and point the client at that.

Browser JavaScript can read X-Request-Id, X-Sample-Rate, X-Output-Format, the X-RateLimit-* headers and Retry-After; everything else is hidden by the browser. (X-Stream-Format is CORS-exposed too, but do not write a read for it here — only the native POST /v1/tts/stream sends it. On this route the framing is already unwrapped for you, so the header is always absent.)

Troubleshooting by symptom​

No audio​

Work down this list — each has a different cause.

  • The saved file contains JSON. On the raw streaming path, a generation that produced nothing answers 502 with an error body where audio was expected. A client that writes the response body straight to a .pcm file ends up with JSON in it. Check the first bytes of the file.
  • The file downloads instead of playing. The synchronous response carries Content-Disposition: attachment. Clients that fetch the body are unaffected; one that navigates to the URL gets a download.
  • Audio plays at the wrong pitch or speed. pcm and ulaw are headerless — the bytes carry no sample rate. pcm is 24000 Hz here unless you asked otherwise; ulaw is always 8000. Read X-Sample-Rate instead of assuming.
  • The player refuses an .mp3. If you omitted response_format and the worker serving you cannot encode mp3, VieNeu falls back to WAV bytes rather than failing. Read X-Output-Format to see what you actually got. (Had you asked for mp3 explicitly, you would have got a 503 naming the format instead — an explicit choice is never silently substituted.)

400 invalid_request_error​

The param field in the error body names the culprit for most of these — with one exception, and it is the first cause listed below. The common causes:

  • An unknown body field. VieNeu rejects fields it does not declare rather than ignoring them. The accepted set is exactly: model, input, voice, response_format, sample_rate, speed, instructions, stream_format, emotion, aiRefine, engine. Note there is no stream boolean — a client that sends stream: true (as for chat completions) gets a 400. So does user, language, or any client-specific extension. This is the most common reason a client that works against OpenAI fails against VieNeu: inspect the request body it actually emits.

    Read message, not param, for this one. The rejection comes from the request validator rather than from the handler, so param comes back as the literal string "property" and never the offending field name. The name is in the message: property stream should not exist. Only the first offending field is reported per request, so if your client adds several, fix and retry until it passes.

  • An unrecognised vieneu-… model name. param: "model". Other unknown model names are ignored; this prefix is not.

  • An OpenAI voice name. alloy, echo, fable, onyx, nova, shimmer are not mapped. param: "voice".

  • A cloned voice id. clone_… works only on POST /v1/tts and POST /v1/tts/stream.

  • stream_format with a non-streamable format. Only pcm and ulaw can be streamed, and response_format defaults to mp3 — so setting stream_format alone is always a 400.

  • aac or flac. Real OpenAI formats VieNeu cannot encode; the message says so explicitly.

  • A sample_rate that contradicts the format. opus is always 48000 and ulaw always 8000; a conflicting rate is rejected, not overridden. Valid rates are 8000, 16000, 22050, 24000, 44100, 48000.

  • speed outside 0.25–4.0. Inside that range it is clamped to 0.5–2.0 and never errors; outside it, validation rejects it.

  • A test key over 100 words. vn_test_ keys cap each request at 100 whitespace-separated words. The message names your actual word count.

401 authentication_error​

  • Invalid API key format. — the key must literally begin vn_sk_ or vn_test_. This is checked before any lookup, so placeholder keys some clients ship or require (sk-..., none, ollama) fail here. If the client refuses to save an empty key field, it needs a real VieNeu key.
  • API key required. although you set one — a bearer token containing a dot is treated as a JWT and ignored as an API key. Real VieNeu keys are hex and never contain a dot, so this means something wrapped or replaced your key.
  • Invalid or revoked API key. — a revoked key is indistinguishable from an unknown one.

Keys go in Authorization: Bearer <key> or X-API-Key: <key>; both work. If your client sends both with different values, X-API-Key silently wins.

403 — usually "out of credit"​

Running out of tokens is 403 here, not 402. Nothing in the public API emits 402, so a client that watches for it never learns it has stopped being paid for. Watch 403 instead. Three messages, all typed insufficient_quota:

  • Insufficient tokens. You have <n> tokens remaining but this request costs <m> tokens. — the ordinary end of a grant, and the normal first failure once a trial allowance empties.
  • Your API token package has expired. Please renew your Developer plan. — time, not consumption. Buying more tokens does not fix it; renewing does.
  • API token grant is not active. Check your Developer plan status.

Two other 403s are not about credit, and the type field tells them apart:

  • Your plan does not include engine "v4". Upgrade your Developer plan to use it. — typed invalid_request_error, param: "engine". Raised only when you pinned an engine explicitly and your plan does not cover it; with v4 the only engine since 2026-09-24 you should rarely see it — drop model/engine and the default engine works.
  • This feature is not available yet. — typed authentication_error, from a feature gate rather than from billing.

So do not branch on type alone to detect "out of credit", and do not branch on 403 alone either. Both together are unambiguous.

A daily or weekly cap is 429, not 403 — see below.

429 rate_limit_exceeded​

Read the headers to tell the sources apart:

Headers presentSourceCounted against
Retry-After + X-RateLimit-*VieNeu's throttler on synthesis routes — 300 requests/min by defaultyour API key
none of them, non-JSON bodythe edge proxy — 3000 requests/minyour source IP, pooled with everyone behind it
JSON, message contains Resets at <ISO>your token grant's daily or weekly capyour grant

Back off in all three cases. The middle one is worth knowing about: if your client runs on shared or NAT'd egress, you can trip a limit no amount of tuning on your side will fix.

Do not hard-code 300. It is a deployment setting (PUBLIC_API_SYNTHESIS_RPM), not a constant, and other pages quote other figures because they were written against other deployments. The row's first column is the durable part: read your live budget from X-RateLimit-Limit and X-RateLimit-Remaining on any successful response.

The third row is the one people mistake for the first. A daily or weekly grant cap is a 429 whose message ends Resets at <ISO timestamp> — no amount of slowing down clears it before that time. Running out of tokens outright is 403, never 402.

503 — the fleet, not your request​

Nothing about your request needs changing; retry shortly. Four causes:

  • No worker for the engine you pinned. No V4 TTS worker is available right now. Please try again shortly. (the streaming path says …available for streaming right now.) Raised only for a non-default engine: a request that takes the default engine falls back to the configured worker rather than 503, so dropping model/engine is a valid workaround here.
  • No active voice on that engine, when you omitted voice. No voice is currently available on engine "v4". Pass an explicit voiceId from GET /v1/voices. This one arrives as a 503 but typed invalid_request_error with param: "voice" — a catalogue problem wearing a request-shaped label. It is raised before billing, so nothing was charged.
  • A format you asked for explicitly that the worker cannot encode. Had you left response_format off you would have got WAV bytes instead; an explicit choice is never silently substituted. See No audio.
  • A sample_rate the worker cannot produce, on the streaming path. The message names both what you asked for and what the worker offered.

Every one of these that got as far as billing refunds automatically — you are not charged for audio you did not receive.

The stream stops mid-sentence​

  • With SSE, wait for the terminal event. There is no data: [DONE] sentinel. Exactly one speech.audio.done (audio complete, carries usage) or speech.audio.error (it is not) always arrives. If neither did, the stream was truncated — discard the audio; the request is refunded automatically.
  • On the raw stream path there is no such signal. A truncated stream is indistinguishable from a short one. If you need proof of completeness, use stream_format: "sse" or the native POST /v1/tts/stream.
  • It "hangs", then dumps everything at once. That used to be vieneu-v3 with stream_format: the endpoint accepted it, returned 200, and v3 emitted its chunks in a burst at the end. Since v3 was retired on 2026-09-24 that request is a 400 instead, so if you still see this pattern the buffering is happening in a proxy or client of your own, not on the engine. Pin vieneu-v4 and check what sits between you and the API.
  • A non-streaming call times out on long text. input has no length cap on this endpoint — the request simply blocks for the full synthesis time, and the client's own HTTP timeout gives up first. The ceilings are a 2 MB request body and a 600s proxy read timeout. For long text use the asynchronous POST /api/v1/tts + GET /api/v1/tts/{jobId} path instead.

Every response carries X-Request-Id. Quote it in a support request.

See also​