# Drop-in for OpenAI-compatible apps

Source: https://docs.vieneu.io/docs/integrations/openai-clients

A lot of software already speaks `POST /v1/audio/speech` and lets you point it at
a custom base URL. VieNeu implements that endpoint, so those apps can use
Vietnamese voices with no plugin and no adapter.

This page is the setup sheet: what to put in which field, for each client.

The VieNeu side of every section below is exact. The client side is written as
"which value goes in which kind of field", because setting labels move between
versions — **verify the field names against your build**.

## The three values

Every client needs the same three things.

| Value | What to enter |
|---|---|
| Base URL | `https://api.vieneu.io/api/v1` (see the trap below) |
| API key | your VieNeu key — starts `vn_sk_` (live) or `vn_test_` (test) |
| Model | `tts-1`, `tts-1-hd`, `gpt-4o-mini-tts`, `vieneu` or `vieneu-v4` → `v4`, the only engine since `v3` was retired on 2026-09-24. Any other name is accepted and ignored — **except** a `vieneu-…` name that is not a live engine, which is a 400. `vieneu-v3` is now one of those. |

The one carve-out in that last row is the one a VieNeu user is most likely to
trip. `vieneu-v5` or `vieneu-turbo` does **not** fall through to the default
engine the way `whatever-tts` does; it returns 400 `invalid_request_error` with
`param: "model"`, listing the engine names that do exist. The prefix is read as
an explicit engine choice, and quietly rendering it on some other engine would
bill at a rate you never picked. If you send both `engine` and `model`, `engine`
wins.

`vieneu-v3` is the case you are most likely to actually hit. It was a valid
choice until `v3` was retired from the cloud API on 2026-09-24, and now answers
400 with `Model 'vieneu-v3' is retired. Engine "v3" was retired on 2026-09-24.
Use engine "v4" and a voice from GET /v1/voices?engine=v4 — the two catalogues
share no ids.` Change the model to `vieneu-v4` (or `tts-1`) and re-check the
voice in the same pass — a `vieneu-…` slug saved from the old `v3` catalogue
will 400 on `param: "voice"` next.

A fourth field, `voice`, is optional but usually wanted: omitting it uses the
first active voice on the resolved engine. What you cannot do is reuse OpenAI's
names — `alloy` / `nova` / etc. are **not** mapped and return 400. Get real ids
from `GET /api/v1/audio/voices` (below).

### The base-URL trap

The path is `/api/v1/audio/speech`. The doubled-looking `/api/v1` is real — the
server sets a global `api` prefix *and* mounts the public API at `v1`. So the
value you type depends on what your client appends to it:

| If the field means… | Enter |
|---|---|
| "the OpenAI base — I append `/audio/speech`" | `https://api.vieneu.io/api/v1` |
| "the host — I append `/v1/audio/speech`" | `https://api.vieneu.io/api` |
| "the full endpoint URL" | `https://api.vieneu.io/api/v1/audio/speech` |

Most clients mean the first. A wrong pick is always a **404, and always JSON** —
but in VieNeu's platform error shape rather than the OpenAI envelope this
endpoint otherwise uses:

```json
{ "statusCode": 404, "message": "Cannot POST /api/v1/v1/audio/speech", "error": "Not Found" }
```

Two things to read off it:

- The `Cannot POST …` message and the **absence** of the
  `{"error":{"message","type","param","code"}}` wrapper that every real
  `/v1/audio/speech` failure carries mean the URL shape is wrong, not the key.
  Do not go hunting your API key on this one.
- The path inside the message is the URL your client actually built. Compare it
  to `/api/v1/audio/speech` and the difference tells you which row above you
  needed — the example here doubled `/v1`, so that client wanted
  `https://api.vieneu.io/api`.

### There is no `/v1/models`

VieNeu serves TTS only. There is no `/v1/models` route and no
`/v1/chat/completions`. Two consequences:

- A **"Test connection" / "Verify" button that probes `/models` will report
  failure** even though synthesis works. Ignore it and send a real request.
- A **model dropdown fed from `/models` will be empty.** Type the model name by
  hand.

In Open WebUI, SillyTavern and LobeChat, this base URL belongs **only** in the
audio/TTS provider setting — never in the general OpenAI/LLM endpoint setting, or
the chat model breaks.

### Prove it works first

Before touching any client, confirm the base URL and the key with curl. Send
`$VIENEU_API_KEY` — your own `vn_sk_…` or `vn_test_…` key — and **no `voice`**,
so nothing but those two values can fail:

```bash
curl https://api.vieneu.io/api/v1/audio/speech \
  -H "Authorization: Bearer $VIENEU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu." }' \
  --output speech.mp3
```

Omitting `voice` makes the server pick the first active voice on the resolved
engine, which is what keeps this step honest: a voice id you have not verified
yet would 400 on `param: "voice"` and leave you unable to tell a bad voice from
a bad key or a bad base URL. Pick a real voice in the next step, once this one
has passed.

If that produces playable audio, everything after this is client configuration.

### Getting voice ids

```bash
# Which engine will my requests land on? No API key needed.
curl https://api.vieneu.io/api/v1/engines

# Voice ids for that engine. API key IS required here.
curl "https://api.vieneu.io/api/v1/audio/voices?engine=v4" \
  -H "Authorization: Bearer $VIENEU_API_KEY"
```

`GET /api/v1/engines` returns one entry per enabled engine with `key`,
`isDefault`, `billingMultiplier` and `features`. Use the one with
`isDefault: true` — since `v3` was retired on 2026-09-24 that is the only entry,
`v4`.

:::warning Voice ids saved before 2026-09-24 may be dead
`audio/voices` now lists `v4` only, filtered or not, and `?engine=v3` is a 400.
But the retired V3 catalogue's ids were `vieneu-…` slugs, V4 ids are display
names (`Ngọc Lan`), and the two shared almost no id space — so a voice a client
saved from an unfiltered list before the retirement will 400 on
`param: "voice"` now. Refresh the picker from the call above, and keep passing
`?engine=v4`: it costs nothing and keeps the request explicit.
:::

Voice matching is case-insensitive but **not** diacritic-insensitive, and V4 ids
contain spaces. A client that slugifies or strips accents before sending will
400 on every request.

Now re-run the smoke test with an id from that list, copied exactly as returned:

```bash
curl https://api.vieneu.io/api/v1/audio/speech \
  -H "Authorization: Bearer $VIENEU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "tts-1", "input": "Xin chào, đây là VieNeu.", "voice": "PASTE_AN_ID_HERE" }' \
  --output speech.mp3
```

If the first curl worked and this one 400s on `param: "voice"`, the id is the
only thing that changed — check the engine filter and the accents before
anything else.

:::note Cloned voices do not work here
`clone_…` ids are rejected on `/v1/audio/speech` with a message naming the two
routes that accept them: `POST /v1/tts` and `POST /v1/tts/stream`. No
OpenAI-compatible client can reach a cloned voice.
:::

## Open WebUI

Audio/TTS settings, OpenAI engine:

| Field | Value |
|---|---|
| TTS engine / provider | the OpenAI-compatible option |
| API base URL | `https://api.vieneu.io/api/v1` |
| API key | `vn_sk_…` |
| Model | `tts-1` |
| Voice | a real id from `audio/voices`, e.g. `Ngọc Lan` |

Notes:

- Put this in the **audio** settings, not the model/connection settings.
- The voice field must accept free text. If your build only offers OpenAI's six
  names, they will all 400 — verify on your build.
- Response format: leave it at the default. Open WebUI plays mp3, which is also
  VieNeu's default.

## SillyTavern

TTS extension, OpenAI-compatible provider:

| Field | Value |
|---|---|
| Provider | the OpenAI TTS option |
| API / base URL | `https://api.vieneu.io/api/v1` |
| API key | `vn_sk_…` |
| Model | `tts-1` |
| Voice (per character) | a real id from `audio/voices` |

Notes:

- SillyTavern assigns a voice per character. Every one of them must be a VieNeu
  id — a character left on `alloy` fails while the others work, which reads as a
  flaky integration.
- If the extension only offers a fixed voice dropdown rather than a text field,
  it cannot address VieNeu voices. Verify on your build.
- If your SillyTavern runs its TTS call from the **browser** rather than its own
  server, see [Browser-side clients](#browser-side-clients).

## LobeChat

TTS / audio settings, OpenAI provider:

| Field | Value |
|---|---|
| OpenAI TTS base URL / proxy URL | `https://api.vieneu.io/api/v1` |
| API key | `vn_sk_…` |
| Model | `tts-1` |
| Voice | a real id from `audio/voices` |

Notes:

- LobeChat keeps separate settings for the chat provider and the TTS provider.
  This URL goes in the **TTS** one only.
- LobeChat is commonly deployed so that the browser calls the TTS provider
  directly — see [Browser-side clients](#browser-side-clients) before you debug
  anything else.

## LiteLLM

Add VieNeu as a model in `config.yaml`:

```yaml
model_list:
  - model_name: vieneu-tts
    litellm_params:
      model: openai/tts-1
      api_base: https://api.vieneu.io/api/v1
      api_key: os.environ/VIENEU_API_KEY
```

Then, with the proxy running, call it exactly as you would OpenAI:

```bash
curl http://localhost:4000/v1/audio/speech \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "vieneu-tts", "input": "Xin chào", "voice": "Ngọc Lan" }' \
  --output speech.mp3
```

Notes:

- **`LITELLM_MASTER_KEY` is not your VieNeu key.** It is LiteLLM's own proxy
  credential, whatever you set when you started the proxy — the port
  (`:4000` above) is LiteLLM's too. Your `vn_sk_…` key appears only in
  `config.yaml`, reached through `os.environ/VIENEU_API_KEY`; the proxy is what
  attaches it to the upstream call.
- The `openai/` prefix on `model` tells LiteLLM to pass the request through in
  OpenAI's shape, which is what VieNeu answers. The name after it (`tts-1`) is
  what reaches VieNeu.
- The VieNeu values above are exact; the surrounding LiteLLM keys are its
  standard `model_list` shape — confirm them against your LiteLLM version.
- **Do not enable any request-enrichment that adds body fields.** VieNeu rejects
  unknown fields with 400 rather than ignoring them — see
  [400](#400-invalid_request_error).

## LiveKit Agents

The OpenAI plugin takes a base URL and a key:

```python
from livekit.plugins import openai

tts = openai.TTS(
    model="vieneu-v4",
    voice="Ngọc Lan",
    base_url="https://api.vieneu.io/api/v1",
    api_key="vn_sk_...",
)
```

Notes:

- **Pin `vieneu-v4`, or omit `model`.** Either lands on `v4`, the only engine
  since `v3` was retired on 2026-09-24. `vieneu-v3` — which this endpoint used
  to accept with streaming, answering 200 and delivering v3's chunks in a burst
  at the end with no error telling you so — is now a 400 with `param: "model"`
  saying the engine was retired.
- `GET /v1/engines` returns the live billing multiplier; see
  [Engines](../cloud-api/overview#engines).
- For streaming, VieNeu needs **both** `response_format: "pcm"` and
  `stream_format`. `response_format` defaults to mp3, so setting only
  `stream_format` is a 400. Whether the plugin sends either, and whether it lets
  you add them, is the thing to verify on your build. Without them the call
  still works — it just returns a complete mp3 instead of a stream.
- `pcm` is headerless and defaults to **24000 Hz** on this route (the engines'
  native rate is 48000). Read the actual rate from the `X-Sample-Rate` response
  header and configure the pipeline to match, or the audio plays at the wrong
  pitch.

## Pipecat

```python
import os
from pipecat.services.openai.tts import OpenAITTSService

tts = OpenAITTSService(
    api_key=os.environ["VIENEU_API_KEY"],
    base_url="https://api.vieneu.io/api/v1",
    model="vieneu-v4",
    voice="Ngọc Lan",
)
```

Notes:

- The import path for Pipecat's OpenAI TTS service has moved between releases —
  check yours. The constructor arguments and the values above are what matter.
- The same three streaming rules as LiveKit apply: pin `vieneu-v4`, send
  `response_format: "pcm"` **and** `stream_format`, and take the rate from
  `X-Sample-Rate` rather than assuming 48 kHz.
- If your version hard-codes a `response_format` VieNeu does not accept
  (`aac`, `flac`), the request 400s naming the field. `wav`, `mp3`, `opus`,
  `pcm` and `ulaw` are the accepted values.

## Browser-side clients

In production, VieNeu's CORS policy allows only VieNeu's own origins. A client
that calls `/v1/audio/speech` **from the page** rather than from its own server
fails at the preflight, with no response body to explain it.

Whether a given LobeChat or SillyTavern deployment does its TTS server-side is a
per-deployment question — verify on your build. If it is browser-side, put a
proxy of your own in front (LiteLLM works well for this) and point the client at
that.

Browser JavaScript can read `X-Request-Id`, `X-Sample-Rate`, `X-Output-Format`,
the `X-RateLimit-*` headers and `Retry-After`; everything else is hidden by the
browser. (`X-Stream-Format` is CORS-exposed too, but do not write a read for it
here — only the native `POST /v1/tts/stream` sends it. On this route the framing
is already unwrapped for you, so the header is always absent.)

## Troubleshooting by symptom

### No audio

Work down this list — each has a different cause.

- **The saved file contains JSON.** On the raw streaming path, a generation that
  produced nothing answers **502** with an error body where audio was expected. A
  client that writes the response body straight to a `.pcm` file ends up with
  JSON in it. Check the first bytes of the file.
- **The file downloads instead of playing.** The synchronous response carries
  `Content-Disposition: attachment`. Clients that fetch the body are unaffected;
  one that navigates to the URL gets a download.
- **Audio plays at the wrong pitch or speed.** `pcm` and `ulaw` are headerless —
  the bytes carry no sample rate. `pcm` is 24000 Hz here unless you asked
  otherwise; `ulaw` is always 8000. Read `X-Sample-Rate` instead of assuming.
- **The player refuses an `.mp3`.** If you omitted `response_format` and the
  worker serving you cannot encode mp3, VieNeu falls back to **WAV bytes** rather
  than failing. Read `X-Output-Format` to see what you actually got. (Had you
  asked for mp3 explicitly, you would have got a 503 naming the format instead —
  an explicit choice is never silently substituted.)

### 400 `invalid_request_error`

The `param` field in the error body names the culprit for most of these — with
one exception, and it is the first cause listed below. The common causes:

- **An unknown body field.** VieNeu rejects fields it does not declare rather
  than ignoring them. The accepted set is exactly: `model`, `input`, `voice`,
  `response_format`, `sample_rate`, `speed`, `instructions`, `stream_format`,
  `emotion`, `aiRefine`, `engine`. Note there is **no `stream` boolean** — a
  client that sends `stream: true` (as for chat completions) gets a 400. So does
  `user`, `language`, or any client-specific extension. This is the most common
  reason a client that works against OpenAI fails against VieNeu: inspect the
  request body it actually emits.

  **Read `message`, not `param`, for this one.** The rejection comes from the
  request validator rather than from the handler, so `param` comes back as the
  literal string `"property"` and never the offending field name. The name is in
  the message: `property stream should not exist`. Only the first offending
  field is reported per request, so if your client adds several, fix and retry
  until it passes.
- **An unrecognised `vieneu-…` model name.** `param: "model"`. Other unknown
  model names are ignored; this prefix is not.
- **An OpenAI voice name.** `alloy`, `echo`, `fable`, `onyx`, `nova`, `shimmer`
  are not mapped. `param: "voice"`.
- **A cloned voice id.** `clone_…` works only on `POST /v1/tts` and
  `POST /v1/tts/stream`.
- **`stream_format` with a non-streamable format.** Only `pcm` and `ulaw` can be
  streamed, and `response_format` defaults to **mp3** — so setting
  `stream_format` alone is always a 400.
- **`aac` or `flac`.** Real OpenAI formats VieNeu cannot encode; the message says
  so explicitly.
- **A `sample_rate` that contradicts the format.** `opus` is always 48000 and
  `ulaw` always 8000; a conflicting rate is rejected, not overridden. Valid rates
  are 8000, 16000, 22050, 24000, 44100, 48000.
- **`speed` outside 0.25–4.0.** Inside that range it is clamped to 0.5–2.0 and
  never errors; outside it, validation rejects it.
- **A test key over 100 words.** `vn_test_` keys cap each request at 100
  whitespace-separated words. The message names your actual word count.

### 401 `authentication_error`

- **`Invalid API key format.`** — the key must literally begin `vn_sk_` or
  `vn_test_`. This is checked before any lookup, so placeholder keys some clients
  ship or require (`sk-...`, `none`, `ollama`) fail here. If the client refuses
  to save an empty key field, it needs a real VieNeu key.
- **`API key required.`** although you set one — a bearer token containing a dot
  is treated as a JWT and ignored as an API key. Real VieNeu keys are hex and
  never contain a dot, so this means something wrapped or replaced your key.
- **`Invalid or revoked API key.`** — a revoked key is indistinguishable from an
  unknown one.

Keys go in `Authorization: Bearer <key>` or `X-API-Key: <key>`; both work. If
your client sends both with different values, `X-API-Key` silently wins.

### 403 — usually "out of credit" {#403-out-of-credit}

**Running out of tokens is 403 here, not 402.** Nothing in the public API emits
402, so a client that watches for it never learns it has stopped being paid for.
Watch 403 instead. Three messages, all typed `insufficient_quota`:

- `Insufficient tokens. You have <n> tokens remaining but this request costs <m> tokens.`
  — the ordinary end of a grant, and the normal first failure once a trial
  allowance empties.
- `Your API token package has expired. Please renew your Developer plan.` —
  time, not consumption. Buying more tokens does not fix it; renewing does.
- `API token grant is not active. Check your Developer plan status.`

Two other 403s are *not* about credit, and the `type` field tells them apart:

- `Your plan does not include engine "v4". Upgrade your Developer plan to
  use it.` — typed `invalid_request_error`, `param: "engine"`. Raised only when
  you pinned an engine explicitly and your plan does not cover it; with `v4`
  the only engine since 2026-09-24 you should rarely see it — drop
  `model`/`engine` and the default engine works.
- `This feature is not available yet.` — typed `authentication_error`, from a
  feature gate rather than from billing.

So do not branch on `type` alone to detect "out of credit", and do not branch on
403 alone either. Both together are unambiguous.

A daily or weekly cap is **429**, not 403 — see below.

### 429 `rate_limit_exceeded`

Read the headers to tell the sources apart:

| Headers present | Source | Counted against |
|---|---|---|
| `Retry-After` + `X-RateLimit-*` | VieNeu's throttler on synthesis routes — 300 requests/min **by default** | your API key |
| none of them, non-JSON body | the edge proxy — 3000 requests/min | your **source IP**, pooled with everyone behind it |
| JSON, message contains `Resets at <ISO>` | your token grant's daily or weekly cap | your grant |

Back off in all three cases. The middle one is worth knowing about: if your
client runs on shared or NAT'd egress, you can trip a limit no amount of tuning
on your side will fix.

**Do not hard-code 300.** It is a deployment setting
(`PUBLIC_API_SYNTHESIS_RPM`), not a constant, and other pages quote other
figures because they were written against other deployments. The row's first
column is the durable part: read your live budget from `X-RateLimit-Limit` and
`X-RateLimit-Remaining` on any successful response.

The third row is the one people mistake for the first. A daily or weekly grant
cap is a 429 whose message ends `Resets at <ISO timestamp>` — no amount of
slowing down clears it before that time. Running out of tokens outright is
[403](#403-out-of-credit), never 402.

### 503 — the fleet, not your request

Nothing about your request needs changing; retry shortly. Four causes:

- **No worker for the engine you pinned.** `No V4 TTS worker is available right
  now. Please try again shortly.` (the streaming path says `…available for
  streaming right now`.) Raised only for a **non-default** engine: a request
  that takes the default engine falls back to the configured worker rather than
  503, so dropping `model`/`engine` is a valid workaround here.
- **No active voice on that engine**, when you omitted `voice`. `No voice is
  currently available on engine "v4". Pass an explicit voiceId from GET
  /v1/voices.` This one arrives as a 503 but typed `invalid_request_error` with
  `param: "voice"` — a catalogue problem wearing a request-shaped label. It is
  raised before billing, so nothing was charged.
- **A format you asked for explicitly that the worker cannot encode.** Had you
  left `response_format` off you would have got WAV bytes instead; an explicit
  choice is never silently substituted. See [No audio](#no-audio).
- **A `sample_rate` the worker cannot produce**, on the streaming path. The
  message names both what you asked for and what the worker offered.

Every one of these that got as far as billing refunds automatically — you are
not charged for audio you did not receive.

### The stream stops mid-sentence

- **With SSE, wait for the terminal event.** There is no `data: [DONE]`
  sentinel. Exactly one `speech.audio.done` (audio complete, carries `usage`) or
  `speech.audio.error` (it is not) always arrives. If neither did, the stream was
  truncated — discard the audio; the request is refunded automatically.
- **On the raw stream path there is no such signal.** A truncated stream is
  indistinguishable from a short one. If you need proof of completeness, use
  `stream_format: "sse"` or the native
  [`POST /v1/tts/stream`](../cloud-api/streaming).
- **It "hangs", then dumps everything at once.** That used to be `vieneu-v3`
  with `stream_format`: the endpoint accepted it, returned 200, and v3 emitted
  its chunks in a burst at the end. Since `v3` was retired on 2026-09-24 that
  request is a 400 instead, so if you still see this pattern the buffering is
  happening in a proxy or client of your own, not on the engine. Pin
  `vieneu-v4` and check what sits between you and the API.
- **A non-streaming call times out on long text.** `input` has **no length cap**
  on this endpoint — the request simply blocks for the full synthesis time, and
  the client's own HTTP timeout gives up first. The ceilings are a 2 MB request
  body and a 600s proxy read timeout. For long text use the asynchronous
  `POST /api/v1/tts` + `GET /api/v1/tts/{jobId}` path instead.

Every response carries `X-Request-Id`. Quote it in a support request.

## See also

- [Integrations overview](./overview.md) — which route to take if your tool is
  not on this page.
- [OpenAI-compatible TTS endpoint](../cloud-api/openai-compatible) — the full
  field-by-field contract for `/v1/audio/speech`: every accepted field, the
  format and sample-rate rules, and the complete error list.
- [Streaming](../cloud-api/streaming) — both streaming endpoints compared,
  and the terminal-event semantics summarised above.
- [Rate limits and request ids](../cloud-api/overview#rate-limits-and-request-ids)
  — the headers, and why limits count per key.
- [API reference](/api-reference) — generated from the server's own OpenAPI spec.
