# Errors

Source: https://docs.vieneu.io/docs/cloud-api/errors

Statuses tell you how bad it was. Most refusals additionally carry a `code`, and
where one exists it is the only part of the body safe to branch on — messages are
prose, some of it Vietnamese, and all of it subject to rewording.

The important thing to know first is that **`code` is not universal**. Most of
the ways a `/v1` call can be refused carry one; a few give you a status and a
sentence. This page says which is which, because a client written on the
assumption that every error has a `code` will silently fall through to its
default branch on the ones that do not.

## Two body shapes

Native endpoints return the platform's shape:

```json
{ "statusCode": 403, "message": "Insufficient tokens", "traceId": "0f7c…" }
```

`POST /v1/audio/speech` returns OpenAI's shape instead, so an OpenAI client can
parse it without an adapter:

```json
{ "error": { "message": "…", "type": "rate_limit_exceeded", "param": null, "code": null } }
```

That is the whole difference: one route, one envelope. Everything else on `/v1` —
including `/v1/vapi/speech`, which serves a third party — uses the first shape.

Every native body repeats the status as `statusCode` and adds `code` and its
extras where the refusal carries one:

```json
{ "statusCode": 429,
  "message": "Server is at capacity. Please try again in a few minutes.",
  "code": "QUEUE_FULL", "traceId": "0f7c…" }
```

A body that carried a `code` used to arrive without `statusCode` — so gaining a
machine-readable reason cost you a field you had always been able to read. That
gap is closed: the field is on every native body now. The HTTP status line is
still the authoritative copy — read it from the response where you can.

## Refusals that carry a `code`

### Quota and grant

The six routes that bill per character or per second — `POST /v1/tts`,
`/v1/tts/stream`, `/v1/dialogue`, `/v1/dub`, `/v1/srt` and
`/v1/vapi/speech` — answer a quota refusal with a code, and the two that recover
on their own say when:

| `code` | Status | What it means | What to do |
|---|---|---|---|
| `GRANT_EXPIRED` | 403 | The token package behind the key has passed its expiry date. | Renew. **Retrying will not help**, and neither will topping up. |
| `GRANT_TOKENS_EXHAUSTED` | 403 | The grant's balance is smaller than this request's price. Carries `remaining` and `required`. | Top up. Retrying will not help. |
| `GRANT_DAILY_LIMIT` | 429 | The plan's daily cap is spent. Carries `resetAt`, `remaining`, `required`. | Sleep until `resetAt`. |
| `GRANT_WEEKLY_LIMIT` | 429 | The plan's weekly cap is spent. Carries `resetAt`, `remaining`, `required`. | Sleep until `resetAt`. |
| `CONCURRENT_CONFLICT` | 429 | Two of your own requests raced for the same balance and this one lost. Carries `retryAfterSeconds: 1`, also sent as `Retry-After: 1`. | Wait the second it names, then retry — this one clears. |

`resetAt` is ISO 8601, and it replaces parsing the timestamp out of the message.
Take it over a guessed backoff, and over an "upgrade your plan" prompt for
something that resets in an hour.

The 403/429 split is the coarse version of the same decision: **403 means
retrying will not help** even though it looks transient, and 429 means it will,
eventually. What the status cannot always tell you is *when*. The daily and
weekly caps carry **no `Retry-After`** — they carry `resetAt` instead, which is
the honest answer for a wait measured in hours. The refusals that clear in
seconds do carry it: `CONCURRENT_CONFLICT` (`Retry-After: 1`) and the capacity
codes below (`Retry-After: 15`). `code` is what separates a race you retry from a
cap you wait out, and `resetAt` is what says how long the wait is.

`FREE_DAILY_LIMIT` and `GRANT_INACTIVE` appear in the platform's internal code
list but not here: an API key always bills a grant, so the free tier is
unreachable, and a grant that is not active answers 403 with prose and no code.

### Capacity

| `code` | Status | What it means | What to do |
|---|---|---|---|
| `QUEUE_FULL` | 429 | The asynchronous job queue is at its depth limit. Carries `retryAfterSeconds` (`Retry-After`). | Wait what the header says and resubmit — with the same `Idempotency-Key` if you sent one. Nothing was charged. |
| `USER_QUEUE_FULL` | 429 | *Your* account already holds its ceiling of queued jobs (30 by default). Carries `retryAfterSeconds`. | Poll what you have queued, then resubmit. Nothing was charged. |
| `STREAM_CONCURRENCY` | 429 | You already hold the maximum number of open streams. | Close one, or wait for one to finish. Nothing was charged. |
| `STREAM_BUSY` | 503 | Every node is healthy but streaming capacity is used up. | Carries `"fallback": "generate"`, and that is a real instruction — the queued path has separate capacity. See [How a stream ends](./streaming.md#how-a-stream-ends). |

`QUEUE_FULL` and `USER_QUEUE_FULL` reach you from `POST /v1/tts` only; the
synchronous routes do not queue.

### Idempotency

Only on `POST /v1/tts`, and only when you sent an `Idempotency-Key` — see
[Retrying safely](./overview.md#retrying-safely).

| `code` | Status | What it means | What to do |
|---|---|---|---|
| `IDEMPOTENCY_KEY_INVALID` | 400 | The header is not 8–128 printable ASCII characters without whitespace. | Send a UUID. |
| `IDEMPOTENCY_IN_FLIGHT` | 409 | A request with this key is still being processed. Carries `retryAfterSeconds: 1` (`Retry-After: 1`). | Retry after the second, with the **same** key. A new key can charge you twice. |
| `IDEMPOTENCY_KEY_REUSED` | 422 | This key was already used with a different request body. | Use a fresh key for a different request; keys are per request, not per session. |

### Cloning

Cloning moved to the web Studio. The enrolment routes — `POST /v1/voices`,
`DELETE /v1/voices/{voiceId}`, `POST /v1/clone`, `POST /v1/upload` and
`POST /v1/prepare` — answer one refusal now, before any billing or GPU work:

| `code` | Status | What it means | What to do |
|---|---|---|---|
| `CLONE_WEB_ONLY` | 410 | Cloned voices are created at [vieneu.io/#/clone](https://vieneu.io/#/clone), not through the API. Carries `cloneUrl`. | Make the voice in the Studio once. Nothing else in your integration changes: it is then listed by `GET /v1/voices` with your key and its `clone_…` id works as `voiceId` on `POST /v1/tts` and `/v1/tts/stream`. |

410 rather than 404 on purpose: the endpoints existed and were documented, so
the status says the surface moved rather than that you mistyped a path. The
older clone codes (`CLONE_QUOTA_EXCEEDED`, `CLONE_MONTHLY_CAP`,
`CLONE_DAILY_CAP`, `CLONE_REF_*`, `CLONE_TRANSCRIPT_DENSITY`,
`CLONE_TOKENS_INSUFFICIENT` and the rest) belong to those routes and can no
longer reach an API key — the same limits still apply to the Studio, where they
are shown as you clone.

Generating **with** a cloned voice is not part of this and carries no clone
codes. A `clone_…` id that is missing or belongs to someone else fails as an
ordinary `400` with a message and no `code` — see
[Cloned voices](./streaming.md#cloned-voices).

## Refusals that do **not** carry a code

### The OpenAI-compatible route

`POST /v1/audio/speech` carries none of the codes above, despite OpenAI's
envelope having a field for it: its `code` is always `null` on this route. What
you get instead is `type`, which is coarser on purpose. A quota refusal is typed
from its status — `insufficient_quota` on the `403`s, which are the exhausted
and expired grants, and `rate_limit_exceeded` on the `429`s, which are the spent
caps and the races. So "out of credit" and "wait" survive the translation;
`GRANT_EXPIRED` and `GRANT_TOKENS_EXHAUSTED` do not, and neither does the
`resetAt` that would tell you how long to wait. Anything raised before the
handler runs — a bad key, a field the validator rejects — is typed from the
status alone, and that mapping folds `401` and `403` together as
`authentication_error`.

### The public API never returns 402

There is no `402 Payment Required` anywhere on `/v1`. Every money path runs
through the same deduction, and that step answers only `403` or `429`. The spec
has carried a `402` example; the server does not send one. A client that treats
`402` as "out of credit" will read the real `403` as a permissions bug and stop
retrying for the wrong reason.

### Two text validations

Before any tokens are spent, submitted text is checked for two things the model
cannot usefully voice. Both come back as a plain `400` with a message and no
`code`.

| Message begins | Rule |
|---|---|
| `Text appears to be in an unsupported language…` | More than **34%** of the letters are outside the Latin script. Vietnamese diacritics and `đ` are Latin, so this fires on Chinese, Japanese, Korean, Cyrillic and the like — including an otherwise Vietnamese passage carrying a long non-Latin quotation. |
| `Text does not look like readable words…` | Either one unbroken run of **more than 30 letters**, or **more than 60%** of the word-like tokens (four letters or longer, and at least three of them present) carry no vowel. |

Punctuation breaks a letter run, so a URL, an email address or a file path
measures as its parts and passes; a keyboard mash still measures as one run and
does not.

**These fire on three routes only:** `POST /v1/tts`, `POST /v1/dialogue` and
`POST /v1/srt`. `POST /v1/tts/stream`, `POST /v1/audio/speech`,
`POST /v1/vapi/speech` and `POST /v1/dub` do not run the check — so text one
route rejects, another will synthesize and bill you for. Worth knowing before you
treat a `400` from one route as proof the text is bad everywhere.

## Stream concurrency

`POST /v1/tts/stream` refuses a new stream while you already hold too many:

| Counted against | Ceiling | On exceeding |
|---|---|---|
| The token grant behind your API key | 4 | `429`, `code: STREAM_CONCURRENCY` |
| A signed-in user in the web app | 2 | `429`, same code |

The slot is taken **before** the balance is touched, so a refused stream leaves
no mark on your tokens. Two details about the counting: it is per backend
process, and the key is the **token grant**, not the API key — so several keys
issued against one grant share the four rather than each getting four of their
own.

## `X-Request-Id`

Every response carries `X-Request-Id`. It is the id our logs are keyed by: one
value spans the API request, the worker call it made, and every log line either
produced.

Most error bodies repeat the value as `traceId`, but **the header is the copy to
read** — it is the only one that is always there. `traceId` is stamped by the
exception filter, so a body written straight to the response never gets one:

- **`POST /v1/audio/speech`**, on every error. The handler writes OpenAI's
  envelope itself, and the filter that would add `traceId` never runs — nor does
  OpenAI's shape have a field for it.
- **`POST /v1/vapi/speech`**, on the `502` it sends when the worker produced no
  audio. That body is written mid-response and carries `statusCode` and
  `message` only.

Quote it in a support request. Without it, "a request failed around 3pm" is a
search; with it, it is a lookup.

The value is always ours. The edge proxy sets `X-Request-ID` on every request it
forwards, unconditionally, so an id you send on the way in is overwritten rather
than adopted — log the one that comes back next to your own correlation id
instead of expecting yours to survive.

## Status summary

| Status | Meaning |
|---|---|
| 400 | Malformed request — an unknown voice, a bad format, text the validator refused, a clone rejection |
| 401 | Missing, malformed or revoked API key |
| 403 | Out of tokens, grant expired, or your plan does not include this engine or feature |
| 413 | Uploaded file over 10 MB |
| 409 | An `Idempotency-Key` whose first call is still running — retry with the same key |
| 422 | Content refused by moderation (only when `aiRefine` is on), or an `Idempotency-Key` reused with a different body |
| 429 | Rate limit, token quota, queue depth or stream concurrency |
| 500 | Synthesis failed. The charge is refunded automatically |
| 502 | Every worker for this engine failed. Refunded |
| 503 | No worker for the requested engine or format, or streaming capacity is full — retry shortly |

`429` is the one status with unrelated causes behind it, and the headers say
which — see [Rate limits](./overview.md#rate-limits-and-request-ids).
