Skip to main content

Errors

Statuses tell you how bad it was. Most refusals additionally carry a code, and where one exists it is the only part of the body safe to branch on — messages are prose, some of it Vietnamese, and all of it subject to rewording.

The important thing to know first is that code is not universal. Most of the ways a /v1 call can be refused carry one; a few give you a status and a sentence. This page says which is which, because a client written on the assumption that every error has a code will silently fall through to its default branch on the ones that do not.

Two body shapes​

Native endpoints return the platform's shape:

{ "statusCode": 403, "message": "Insufficient tokens", "traceId": "0f7c…" }

POST /v1/audio/speech returns OpenAI's shape instead, so an OpenAI client can parse it without an adapter:

{ "error": { "message": "…", "type": "rate_limit_exceeded", "param": null, "code": null } }

That is the whole difference: one route, one envelope. Everything else on /v1 — including /v1/vapi/speech, which serves a third party — uses the first shape.

Every native body repeats the status as statusCode and adds code and its extras where the refusal carries one:

{ "statusCode": 429,
"message": "Server is at capacity. Please try again in a few minutes.",
"code": "QUEUE_FULL", "traceId": "0f7c…" }

A body that carried a code used to arrive without statusCode — so gaining a machine-readable reason cost you a field you had always been able to read. That gap is closed: the field is on every native body now. The HTTP status line is still the authoritative copy — read it from the response where you can.

Refusals that carry a code​

Quota and grant​

The six routes that bill per character or per second — POST /v1/tts, /v1/tts/stream, /v1/dialogue, /v1/dub, /v1/srt and /v1/vapi/speech — answer a quota refusal with a code, and the two that recover on their own say when:

codeStatusWhat it meansWhat to do
GRANT_EXPIRED403The token package behind the key has passed its expiry date.Renew. Retrying will not help, and neither will topping up.
GRANT_TOKENS_EXHAUSTED403The grant's balance is smaller than this request's price. Carries remaining and required.Top up. Retrying will not help.
GRANT_DAILY_LIMIT429The plan's daily cap is spent. Carries resetAt, remaining, required.Sleep until resetAt.
GRANT_WEEKLY_LIMIT429The plan's weekly cap is spent. Carries resetAt, remaining, required.Sleep until resetAt.
CONCURRENT_CONFLICT429Two of your own requests raced for the same balance and this one lost. Carries retryAfterSeconds: 1, also sent as Retry-After: 1.Wait the second it names, then retry — this one clears.

resetAt is ISO 8601, and it replaces parsing the timestamp out of the message. Take it over a guessed backoff, and over an "upgrade your plan" prompt for something that resets in an hour.

The 403/429 split is the coarse version of the same decision: 403 means retrying will not help even though it looks transient, and 429 means it will, eventually. What the status cannot always tell you is when. The daily and weekly caps carry no Retry-After — they carry resetAt instead, which is the honest answer for a wait measured in hours. The refusals that clear in seconds do carry it: CONCURRENT_CONFLICT (Retry-After: 1) and the capacity codes below (Retry-After: 15). code is what separates a race you retry from a cap you wait out, and resetAt is what says how long the wait is.

FREE_DAILY_LIMIT and GRANT_INACTIVE appear in the platform's internal code list but not here: an API key always bills a grant, so the free tier is unreachable, and a grant that is not active answers 403 with prose and no code.

Capacity​

codeStatusWhat it meansWhat to do
QUEUE_FULL429The asynchronous job queue is at its depth limit. Carries retryAfterSeconds (Retry-After).Wait what the header says and resubmit — with the same Idempotency-Key if you sent one. Nothing was charged.
USER_QUEUE_FULL429Your account already holds its ceiling of queued jobs (30 by default). Carries retryAfterSeconds.Poll what you have queued, then resubmit. Nothing was charged.
STREAM_CONCURRENCY429You already hold the maximum number of open streams.Close one, or wait for one to finish. Nothing was charged.
STREAM_BUSY503Every node is healthy but streaming capacity is used up.Carries "fallback": "generate", and that is a real instruction — the queued path has separate capacity. See How a stream ends.

QUEUE_FULL and USER_QUEUE_FULL reach you from POST /v1/tts only; the synchronous routes do not queue.

Idempotency​

Only on POST /v1/tts, and only when you sent an Idempotency-Key — see Retrying safely.

codeStatusWhat it meansWhat to do
IDEMPOTENCY_KEY_INVALID400The header is not 8–128 printable ASCII characters without whitespace.Send a UUID.
IDEMPOTENCY_IN_FLIGHT409A request with this key is still being processed. Carries retryAfterSeconds: 1 (Retry-After: 1).Retry after the second, with the same key. A new key can charge you twice.
IDEMPOTENCY_KEY_REUSED422This key was already used with a different request body.Use a fresh key for a different request; keys are per request, not per session.

Cloning​

Cloning moved to the web Studio. The enrolment routes — POST /v1/voices, DELETE /v1/voices/{voiceId}, POST /v1/clone, POST /v1/upload and POST /v1/prepare — answer one refusal now, before any billing or GPU work:

codeStatusWhat it meansWhat to do
CLONE_WEB_ONLY410Cloned voices are created at vieneu.io/#/clone, not through the API. Carries cloneUrl.Make the voice in the Studio once. Nothing else in your integration changes: it is then listed by GET /v1/voices with your key and its clone_… id works as voiceId on POST /v1/tts and /v1/tts/stream.

410 rather than 404 on purpose: the endpoints existed and were documented, so the status says the surface moved rather than that you mistyped a path. The older clone codes (CLONE_QUOTA_EXCEEDED, CLONE_MONTHLY_CAP, CLONE_DAILY_CAP, CLONE_REF_*, CLONE_TRANSCRIPT_DENSITY, CLONE_TOKENS_INSUFFICIENT and the rest) belong to those routes and can no longer reach an API key — the same limits still apply to the Studio, where they are shown as you clone.

Generating with a cloned voice is not part of this and carries no clone codes. A clone_… id that is missing or belongs to someone else fails as an ordinary 400 with a message and no code — see Cloned voices.

Refusals that do not carry a code​

The OpenAI-compatible route​

POST /v1/audio/speech carries none of the codes above, despite OpenAI's envelope having a field for it: its code is always null on this route. What you get instead is type, which is coarser on purpose. A quota refusal is typed from its status — insufficient_quota on the 403s, which are the exhausted and expired grants, and rate_limit_exceeded on the 429s, which are the spent caps and the races. So "out of credit" and "wait" survive the translation; GRANT_EXPIRED and GRANT_TOKENS_EXHAUSTED do not, and neither does the resetAt that would tell you how long to wait. Anything raised before the handler runs — a bad key, a field the validator rejects — is typed from the status alone, and that mapping folds 401 and 403 together as authentication_error.

The public API never returns 402​

There is no 402 Payment Required anywhere on /v1. Every money path runs through the same deduction, and that step answers only 403 or 429. The spec has carried a 402 example; the server does not send one. A client that treats 402 as "out of credit" will read the real 403 as a permissions bug and stop retrying for the wrong reason.

Two text validations​

Before any tokens are spent, submitted text is checked for two things the model cannot usefully voice. Both come back as a plain 400 with a message and no code.

Message beginsRule
Text appears to be in an unsupported language…More than 34% of the letters are outside the Latin script. Vietnamese diacritics and đ are Latin, so this fires on Chinese, Japanese, Korean, Cyrillic and the like — including an otherwise Vietnamese passage carrying a long non-Latin quotation.
Text does not look like readable words…Either one unbroken run of more than 30 letters, or more than 60% of the word-like tokens (four letters or longer, and at least three of them present) carry no vowel.

Punctuation breaks a letter run, so a URL, an email address or a file path measures as its parts and passes; a keyboard mash still measures as one run and does not.

These fire on three routes only: POST /v1/tts, POST /v1/dialogue and POST /v1/srt. POST /v1/tts/stream, POST /v1/audio/speech, POST /v1/vapi/speech and POST /v1/dub do not run the check — so text one route rejects, another will synthesize and bill you for. Worth knowing before you treat a 400 from one route as proof the text is bad everywhere.

Stream concurrency​

POST /v1/tts/stream refuses a new stream while you already hold too many:

Counted againstCeilingOn exceeding
The token grant behind your API key4429, code: STREAM_CONCURRENCY
A signed-in user in the web app2429, same code

The slot is taken before the balance is touched, so a refused stream leaves no mark on your tokens. Two details about the counting: it is per backend process, and the key is the token grant, not the API key — so several keys issued against one grant share the four rather than each getting four of their own.

X-Request-Id​

Every response carries X-Request-Id. It is the id our logs are keyed by: one value spans the API request, the worker call it made, and every log line either produced.

Most error bodies repeat the value as traceId, but the header is the copy to read — it is the only one that is always there. traceId is stamped by the exception filter, so a body written straight to the response never gets one:

  • POST /v1/audio/speech, on every error. The handler writes OpenAI's envelope itself, and the filter that would add traceId never runs — nor does OpenAI's shape have a field for it.
  • POST /v1/vapi/speech, on the 502 it sends when the worker produced no audio. That body is written mid-response and carries statusCode and message only.

Quote it in a support request. Without it, "a request failed around 3pm" is a search; with it, it is a lookup.

The value is always ours. The edge proxy sets X-Request-ID on every request it forwards, unconditionally, so an id you send on the way in is overwritten rather than adopted — log the one that comes back next to your own correlation id instead of expecting yours to survive.

Status summary​

StatusMeaning
400Malformed request — an unknown voice, a bad format, text the validator refused, a clone rejection
401Missing, malformed or revoked API key
403Out of tokens, grant expired, or your plan does not include this engine or feature
413Uploaded file over 10 MB
409An Idempotency-Key whose first call is still running — retry with the same key
422Content refused by moderation (only when aiRefine is on), or an Idempotency-Key reused with a different body
429Rate limit, token quota, queue depth or stream concurrency
500Synthesis failed. The charge is refunded automatically
502Every worker for this engine failed. Refunded
503No worker for the requested engine or format, or streaming capacity is full — retry shortly

429 is the one status with unrelated causes behind it, and the headers say which — see Rate limits.