Drop-in for OpenAI-compatible apps
A lot of software already speaks POST /v1/audio/speech and lets you point it at
a custom base URL. VieNeu implements that endpoint, so those apps can use
Vietnamese voices with no plugin and no adapter.
This page is the setup sheet: what to put in which field, for each client.
The VieNeu side of every section below is exact. The client side is written as "which value goes in which kind of field", because setting labels move between versions — verify the field names against your build.
The three values
Every client needs the same three things.
| Value | What to enter |
|---|---|
| Base URL | https://api.vieneu.io/api/v1 (see the trap below) |
| API key | your VieNeu key — starts vn_sk_ (live) or vn_test_ (test) |
| Model | tts-1, tts-1-hd, gpt-4o-mini-tts, vieneu or vieneu-v4 → v4, the only engine since v3 was retired on 2026-09-24. Any other name is accepted and ignored — except a vieneu-… name that is not a live engine, which is a 400. vieneu-v3 is now one of those. |
The one carve-out in that last row is the one a VieNeu user is most likely to
trip. vieneu-v5 or vieneu-turbo does not fall through to the default
engine the way whatever-tts does; it returns 400 invalid_request_error with
param: "model", listing the engine names that do exist. The prefix is read as
an explicit engine choice, and quietly rendering it on some other engine would
bill at a rate you never picked. If you send both engine and model, engine
wins.
vieneu-v3 is the case you are most likely to actually hit. It was a valid
choice until v3 was retired from the cloud API on 2026-09-24, and now answers
400 with Model 'vieneu-v3' is retired. Engine "v3" was retired on 2026-09-24. Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues share no ids. Change the model to vieneu-v4 (or tts-1) and re-check the
voice in the same pass — a vieneu-… slug saved from the old v3 catalogue
will 400 on param: "voice" next.
A fourth field, voice, is optional but usually wanted: omitting it uses the
first active voice on the resolved engine. What you cannot do is reuse OpenAI's
names — alloy / nova / etc. are not mapped and return 400. Get real ids
from GET /api/v1/audio/voices (below).
The base-URL trap
The path is /api/v1/audio/speech. The doubled-looking /api/v1 is real — the
server sets a global api prefix and mounts the public API at v1. So the
value you type depends on what your client appends to it:
| If the field means… | Enter |
|---|---|
"the OpenAI base — I append /audio/speech" | https://api.vieneu.io/api/v1 |
"the host — I append /v1/audio/speech" | https://api.vieneu.io/api |
| "the full endpoint URL" | https://api.vieneu.io/api/v1/audio/speech |
Most clients mean the first. A wrong pick is always a 404, and always JSON — but in VieNeu's platform error shape rather than the OpenAI envelope this endpoint otherwise uses:
{ "statusCode": 404, "message": "Cannot POST /api/v1/v1/audio/speech", "error": "Not Found" }
Two things to read off it:
- The
Cannot POST …message and the absence of the{"error":{"message","type","param","code"}}wrapper that every real/v1/audio/speechfailure carries mean the URL shape is wrong, not the key. Do not go hunting your API key on this one. - The path inside the message is the URL your client actually built. Compare it
to
/api/v1/audio/speechand the difference tells you which row above you needed — the example here doubled/v1, so that client wantedhttps://api.vieneu.io/api.
There is no /v1/models
VieNeu serves TTS only. There is no /v1/models route and no
/v1/chat/completions. Two consequences:
- A "Test connection" / "Verify" button that probes
/modelswill report failure even though synthesis works. Ignore it and send a real request. - A model dropdown fed from
/modelswill be empty. Type the model name by hand.
In Open WebUI, SillyTavern and LobeChat, this base URL belongs only in the audio/TTS provider setting — never in the general OpenAI/LLM endpoint setting, or the chat model breaks.
Prove it works first
Before touching any client, confirm the base URL and the key with curl. Send
$VIENEU_API_KEY — your own vn_sk_… or vn_test_… key — and no voice,
so nothing but those two values can fail:
curl https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu." }' \
--output speech.mp3
Omitting voice makes the server pick the first active voice on the resolved
engine, which is what keeps this step honest: a voice id you have not verified
yet would 400 on param: "voice" and leave you unable to tell a bad voice from
a bad key or a bad base URL. Pick a real voice in the next step, once this one
has passed.
If that produces playable audio, everything after this is client configuration.
Getting voice ids
# Which engine will my requests land on? No API key needed.
curl https://api.vieneu.io/api/v1/engines
# Voice ids for that engine. API key IS required here.
curl "https://api.vieneu.io/api/v1/audio/voices?engine=v4" \
-H "Authorization: Bearer $VIENEU_API_KEY"
GET /api/v1/engines returns one entry per enabled engine with key,
isDefault, billingMultiplier and features. Use the one with
isDefault: true — since v3 was retired on 2026-09-24 that is the only entry,
v4.
audio/voices now lists v4 only, filtered or not, and ?engine=v3 is a 400.
But the retired V3 catalogue's ids were vieneu-… slugs, V4 ids are display
names (Ngọc Lan), and the two shared almost no id space — so a voice a client
saved from an unfiltered list before the retirement will 400 on
param: "voice" now. Refresh the picker from the call above, and keep passing
?engine=v4: it costs nothing and keeps the request explicit.
Voice matching is case-insensitive but not diacritic-insensitive, and V4 ids contain spaces. A client that slugifies or strips accents before sending will 400 on every request.
Now re-run the smoke test with an id from that list, copied exactly as returned:
curl https://api.vieneu.io/api/v1/audio/speech \
-H "Authorization: Bearer $VIENEU_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu.", "voice": "PASTE_AN_ID_HERE" }' \
--output speech.mp3
If the first curl worked and this one 400s on param: "voice", the id is the
only thing that changed — check the engine filter and the accents before
anything else.
clone_… ids are rejected on /v1/audio/speech with a message naming the two
routes that accept them: POST /v1/tts and POST /v1/tts/stream. No
OpenAI-compatible client can reach a cloned voice.
Open WebUI
Audio/TTS settings, OpenAI engine:
| Field | Value |
|---|---|
| TTS engine / provider | the OpenAI-compatible option |
| API base URL | https://api.vieneu.io/api/v1 |
| API key | vn_sk_… |
| Model | tts-1 |
| Voice | a real id from audio/voices, e.g. Ngọc Lan |
Notes:
- Put this in the audio settings, not the model/connection settings.
- The voice field must accept free text. If your build only offers OpenAI's six names, they will all 400 — verify on your build.
- Response format: leave it at the default. Open WebUI plays mp3, which is also VieNeu's default.
SillyTavern
TTS extension, OpenAI-compatible provider:
| Field | Value |
|---|---|
| Provider | the OpenAI TTS option |
| API / base URL | https://api.vieneu.io/api/v1 |
| API key | vn_sk_… |
| Model | tts-1 |
| Voice (per character) | a real id from audio/voices |
Notes:
- SillyTavern assigns a voice per character. Every one of them must be a VieNeu
id — a character left on
alloyfails while the others work, which reads as a flaky integration. - If the extension only offers a fixed voice dropdown rather than a text field, it cannot address VieNeu voices. Verify on your build.
- If your SillyTavern runs its TTS call from the browser rather than its own server, see Browser-side clients.
LobeChat
TTS / audio settings, OpenAI provider:
| Field | Value |
|---|---|
| OpenAI TTS base URL / proxy URL | https://api.vieneu.io/api/v1 |
| API key | vn_sk_… |
| Model | tts-1 |
| Voice | a real id from audio/voices |
Notes:
- LobeChat keeps separate settings for the chat provider and the TTS provider. This URL goes in the TTS one only.
- LobeChat is commonly deployed so that the browser calls the TTS provider directly — see Browser-side clients before you debug anything else.
LiteLLM
Add VieNeu as a model in config.yaml:
model_list:
- model_name: vieneu-tts
litellm_params:
model: openai/tts-1
api_base: https://api.vieneu.io/api/v1
api_key: os.environ/VIENEU_API_KEY
Then, with the proxy running, call it exactly as you would OpenAI:
curl http://localhost:4000/v1/audio/speech \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "vieneu-tts", "input": "Xin chào", "voice": "Ngọc Lan" }' \
--output speech.mp3
Notes:
LITELLM_MASTER_KEYis not your VieNeu key. It is LiteLLM's own proxy credential, whatever you set when you started the proxy — the port (:4000above) is LiteLLM's too. Yourvn_sk_…key appears only inconfig.yaml, reached throughos.environ/VIENEU_API_KEY; the proxy is what attaches it to the upstream call.- The
openai/prefix onmodeltells LiteLLM to pass the request through in OpenAI's shape, which is what VieNeu answers. The name after it (tts-1) is what reaches VieNeu. - The VieNeu values above are exact; the surrounding LiteLLM keys are its
standard
model_listshape — confirm them against your LiteLLM version. - Do not enable any request-enrichment that adds body fields. VieNeu rejects unknown fields with 400 rather than ignoring them — see 400.
LiveKit Agents
The OpenAI plugin takes a base URL and a key:
from livekit.plugins import openai
tts = openai.TTS(
model="vieneu-v4",
voice="Ngọc Lan",
base_url="https://api.vieneu.io/api/v1",
api_key="vn_sk_...",
)
Notes:
- Pin
vieneu-v4, or omitmodel. Either lands onv4, the only engine sincev3was retired on 2026-09-24.vieneu-v3— which this endpoint used to accept with streaming, answering 200 and delivering v3's chunks in a burst at the end with no error telling you so — is now a 400 withparam: "model"saying the engine was retired. GET /v1/enginesreturns the live billing multiplier; see Engines.- For streaming, VieNeu needs both
response_format: "pcm"andstream_format.response_formatdefaults to mp3, so setting onlystream_formatis a 400. Whether the plugin sends either, and whether it lets you add them, is the thing to verify on your build. Without them the call still works — it just returns a complete mp3 instead of a stream. pcmis headerless and defaults to 24000 Hz on this route (the engines' native rate is 48000). Read the actual rate from theX-Sample-Rateresponse header and configure the pipeline to match, or the audio plays at the wrong pitch.
Pipecat
import os
from pipecat.services.openai.tts import OpenAITTSService
tts = OpenAITTSService(
api_key=os.environ["VIENEU_API_KEY"],
base_url="https://api.vieneu.io/api/v1",
model="vieneu-v4",
voice="Ngọc Lan",
)
Notes:
- The import path for Pipecat's OpenAI TTS service has moved between releases — check yours. The constructor arguments and the values above are what matter.
- The same three streaming rules as LiveKit apply: pin
vieneu-v4, sendresponse_format: "pcm"andstream_format, and take the rate fromX-Sample-Raterather than assuming 48 kHz. - If your version hard-codes a
response_formatVieNeu does not accept (aac,flac), the request 400s naming the field.wav,mp3,opus,pcmandulaware the accepted values.
Browser-side clients
In production, VieNeu's CORS policy allows only VieNeu's own origins. A client
that calls /v1/audio/speech from the page rather than from its own server
fails at the preflight, with no response body to explain it.
Whether a given LobeChat or SillyTavern deployment does its TTS server-side is a per-deployment question — verify on your build. If it is browser-side, put a proxy of your own in front (LiteLLM works well for this) and point the client at that.
Browser JavaScript can read X-Request-Id, X-Sample-Rate, X-Output-Format,
the X-RateLimit-* headers and Retry-After; everything else is hidden by the
browser. (X-Stream-Format is CORS-exposed too, but do not write a read for it
here — only the native POST /v1/tts/stream sends it. On this route the framing
is already unwrapped for you, so the header is always absent.)
Troubleshooting by symptom
No audio
Work down this list — each has a different cause.
- The saved file contains JSON. On the raw streaming path, a generation that
produced nothing answers 502 with an error body where audio was expected. A
client that writes the response body straight to a
.pcmfile ends up with JSON in it. Check the first bytes of the file. - The file downloads instead of playing. The synchronous response carries
Content-Disposition: attachment. Clients that fetch the body are unaffected; one that navigates to the URL gets a download. - Audio plays at the wrong pitch or speed.
pcmandulaware headerless — the bytes carry no sample rate.pcmis 24000 Hz here unless you asked otherwise;ulawis always 8000. ReadX-Sample-Rateinstead of assuming. - The player refuses an
.mp3. If you omittedresponse_formatand the worker serving you cannot encode mp3, VieNeu falls back to WAV bytes rather than failing. ReadX-Output-Formatto see what you actually got. (Had you asked for mp3 explicitly, you would have got a 503 naming the format instead — an explicit choice is never silently substituted.)
400 invalid_request_error
The param field in the error body names the culprit for most of these — with
one exception, and it is the first cause listed below. The common causes:
-
An unknown body field. VieNeu rejects fields it does not declare rather than ignoring them. The accepted set is exactly:
model,input,voice,response_format,sample_rate,speed,instructions,stream_format,emotion,aiRefine,engine. Note there is nostreamboolean — a client that sendsstream: true(as for chat completions) gets a 400. So doesuser,language, or any client-specific extension. This is the most common reason a client that works against OpenAI fails against VieNeu: inspect the request body it actually emits.Read
message, notparam, for this one. The rejection comes from the request validator rather than from the handler, soparamcomes back as the literal string"property"and never the offending field name. The name is in the message:property stream should not exist. Only the first offending field is reported per request, so if your client adds several, fix and retry until it passes. -
An unrecognised
vieneu-…model name.param: "model". Other unknown model names are ignored; this prefix is not. -
An OpenAI voice name.
alloy,echo,fable,onyx,nova,shimmerare not mapped.param: "voice". -
A cloned voice id.
clone_…works only onPOST /v1/ttsandPOST /v1/tts/stream. -
stream_formatwith a non-streamable format. Onlypcmandulawcan be streamed, andresponse_formatdefaults to mp3 — so settingstream_formatalone is always a 400. -
aacorflac. Real OpenAI formats VieNeu cannot encode; the message says so explicitly. -
A
sample_ratethat contradicts the format.opusis always 48000 andulawalways 8000; a conflicting rate is rejected, not overridden. Valid rates are 8000, 16000, 22050, 24000, 44100, 48000. -
speedoutside 0.25–4.0. Inside that range it is clamped to 0.5–2.0 and never errors; outside it, validation rejects it. -
A test key over 100 words.
vn_test_keys cap each request at 100 whitespace-separated words. The message names your actual word count.
401 authentication_error
Invalid API key format.— the key must literally beginvn_sk_orvn_test_. This is checked before any lookup, so placeholder keys some clients ship or require (sk-...,none,ollama) fail here. If the client refuses to save an empty key field, it needs a real VieNeu key.API key required.although you set one — a bearer token containing a dot is treated as a JWT and ignored as an API key. Real VieNeu keys are hex and never contain a dot, so this means something wrapped or replaced your key.Invalid or revoked API key.— a revoked key is indistinguishable from an unknown one.
Keys go in Authorization: Bearer <key> or X-API-Key: <key>; both work. If
your client sends both with different values, X-API-Key silently wins.
403 — usually "out of credit"
Running out of tokens is 403 here, not 402. Nothing in the public API emits
402, so a client that watches for it never learns it has stopped being paid for.
Watch 403 instead. Three messages, all typed insufficient_quota:
Insufficient tokens. You have <n> tokens remaining but this request costs <m> tokens.— the ordinary end of a grant, and the normal first failure once a trial allowance empties.Your API token package has expired. Please renew your Developer plan.— time, not consumption. Buying more tokens does not fix it; renewing does.API token grant is not active. Check your Developer plan status.
Two other 403s are not about credit, and the type field tells them apart:
Your plan does not include engine "v4". Upgrade your Developer plan to use it.— typedinvalid_request_error,param: "engine". Raised only when you pinned an engine explicitly and your plan does not cover it; withv4the only engine since 2026-09-24 you should rarely see it — dropmodel/engineand the default engine works.This feature is not available yet.— typedauthentication_error, from a feature gate rather than from billing.
So do not branch on type alone to detect "out of credit", and do not branch on
403 alone either. Both together are unambiguous.
A daily or weekly cap is 429, not 403 — see below.
429 rate_limit_exceeded
Read the headers to tell the sources apart:
| Headers present | Source | Counted against |
|---|---|---|
Retry-After + X-RateLimit-* | VieNeu's throttler on synthesis routes — 300 requests/min by default | your API key |
| none of them, non-JSON body | the edge proxy — 3000 requests/min | your source IP, pooled with everyone behind it |
JSON, message contains Resets at <ISO> | your token grant's daily or weekly cap | your grant |
Back off in all three cases. The middle one is worth knowing about: if your client runs on shared or NAT'd egress, you can trip a limit no amount of tuning on your side will fix.
Do not hard-code 300. It is a deployment setting
(PUBLIC_API_SYNTHESIS_RPM), not a constant, and other pages quote other
figures because they were written against other deployments. The row's first
column is the durable part: read your live budget from X-RateLimit-Limit and
X-RateLimit-Remaining on any successful response.
The third row is the one people mistake for the first. A daily or weekly grant
cap is a 429 whose message ends Resets at <ISO timestamp> — no amount of
slowing down clears it before that time. Running out of tokens outright is
403, never 402.
503 — the fleet, not your request
Nothing about your request needs changing; retry shortly. Four causes:
- No worker for the engine you pinned.
No V4 TTS worker is available right now. Please try again shortly.(the streaming path says…available for streaming right now.) Raised only for a non-default engine: a request that takes the default engine falls back to the configured worker rather than 503, so droppingmodel/engineis a valid workaround here. - No active voice on that engine, when you omitted
voice.No voice is currently available on engine "v4". Pass an explicit voiceId from GET /v1/voices.This one arrives as a 503 but typedinvalid_request_errorwithparam: "voice"— a catalogue problem wearing a request-shaped label. It is raised before billing, so nothing was charged. - A format you asked for explicitly that the worker cannot encode. Had you
left
response_formatoff you would have got WAV bytes instead; an explicit choice is never silently substituted. See No audio. - A
sample_ratethe worker cannot produce, on the streaming path. The message names both what you asked for and what the worker offered.
Every one of these that got as far as billing refunds automatically — you are not charged for audio you did not receive.
The stream stops mid-sentence
- With SSE, wait for the terminal event. There is no
data: [DONE]sentinel. Exactly onespeech.audio.done(audio complete, carriesusage) orspeech.audio.error(it is not) always arrives. If neither did, the stream was truncated — discard the audio; the request is refunded automatically. - On the raw stream path there is no such signal. A truncated stream is
indistinguishable from a short one. If you need proof of completeness, use
stream_format: "sse"or the nativePOST /v1/tts/stream. - It "hangs", then dumps everything at once. That used to be
vieneu-v3withstream_format: the endpoint accepted it, returned 200, and v3 emitted its chunks in a burst at the end. Sincev3was retired on 2026-09-24 that request is a 400 instead, so if you still see this pattern the buffering is happening in a proxy or client of your own, not on the engine. Pinvieneu-v4and check what sits between you and the API. - A non-streaming call times out on long text.
inputhas no length cap on this endpoint — the request simply blocks for the full synthesis time, and the client's own HTTP timeout gives up first. The ceilings are a 2 MB request body and a 600s proxy read timeout. For long text use the asynchronousPOST /api/v1/tts+GET /api/v1/tts/{jobId}path instead.
Every response carries X-Request-Id. Quote it in a support request.
See also
- Integrations overview — which route to take if your tool is not on this page.
- OpenAI-compatible TTS endpoint — the full
field-by-field contract for
/v1/audio/speech: every accepted field, the format and sample-rate rules, and the complete error list. - Streaming — both streaming endpoints compared, and the terminal-event semantics summarised above.
- Rate limits and request ids — the headers, and why limits count per key.
- API reference — generated from the server's own OpenAPI spec.